megachangelog
Announcement

AGENTS.md outperforms skills in agent evals

Vercel's testing shows that embedding a compressed 8KB documentation index directly in AGENTS.md achieves 100% pass rate on Next.js 16 agent tasks, significantly outperforming the skill-based approach which maxed out at 79%. This finding challenges the assumption that agent-invoked tools are superior to passive context injection for code generation tasks.

We expected to be the solution for teaching coding agents framework-specific knowledge. After building evals focused on Next.js 16 APIs, we found something unexpected.skills

A compressed 8KB docs index embedded directly in achieved a 100% pass rate, while skills maxed out at 79% even with explicit instructions telling the agent to use them. Without those instructions, skills performed no better than having no documentation at all.AGENTS.md

Here's what we tried, what we learned, and how you can set this up for your own Next.js projects.

AI coding agents rely on training data that becomes outdated. APIs like , , and that aren't in current model training data. When agents don't know these APIs, they generate incorrect code or fall back to older patterns.Next.js 16 introduces'use cache'connection()forbidden()

The reverse can also be true, where you're running an older Next.js version and the model suggests newer APIs that don't exist in your project yet. We wanted to fix this by giving agents access to version-matched documentation.

Before diving into results, a quick explanation of the two approaches we tested:

We built a Next.js docs skill and an docs index, then ran them through our eval suite to see which performed better.AGENTS.md

Skills seemed like the right abstraction. You package your framework docs into a skill, the agent invokes it when working on Next.js tasks, and you get correct code. Clean separation of concerns, minimal context overhead, and the agent only loads what it needs. There's even a growing directory of reusable skills at .skills.sh

We expected the agent to encounter a Next.js task, invoke the skill, read version-matched docs, and generate correct code.

Then we ran the evals.

In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the skill produced no improvement over baseline:

Zero improvement. The skill existed, the agent could use it, and the agent chose not to. On the detailed Build/Lint/Test breakdown, the skill actually performed worse than baseline on some metrics (58% vs 63% on tests), suggesting that an unused skill in the environment may introduce noise or distraction.

This isn't unique to our setup. Agents not reliably using available tools is a of current models.known limitation

We tried adding explicit instructions to telling the agent to use the skill.AGENTS.md

This improved the trigger rate to 95%+ and boosted the pass rate to 79%.

A solid improvement. But we discovered something unexpected about how the instruction wording affected agent behavior.

Different wordings produced dramatically different results:

Same skill. Same docs. Different outcomes based on subtle wording changes.

In one eval (the directive test), the "invoke first" approach wrote correct but completely missed the required changes. The "explore first" approach got both.'use cache'page.tsxnext.config.ts

This fragility concerned us. If small wording tweaks produce large behavioral swings, the approach feels brittle for production use.

Before drawing conclusions, we needed evals we could trust. Our initial test suite had ambiguous prompts, tests that validated implementation details rather than observable behavior, and a focus on APIs already in model training data. We weren't measuring what we actually cared about.

We hardened the eval suite by removing test leakage, resolving contradictions, and shifting to behavior-based assertions. Most importantly, we added tests targeting Next.js 16 APIs that aren't in model training data.

APIs in our focused eval suite:

All the results that follow come from this hardened eval suite. Every configuration was judged against the same tests, with retries to rule out model variance.

What if we removed the decision entirely? Instead of hoping agents would invoke a skill, we could embed a docs index directly in . Not the full documentation, just an index that tells the agent where to find specific doc files that match your project's Next.js version. The agent can then read those files as needed, getting version-accurate information whether you're on the latest release or maintaining an older project.AGENTS.md

We added a key instruction to the injected content.

This tells the agent to consult the docs rather than rely on potentially outdated training data.

We ran the hardened eval suite across all four configurations:

Final pass rates:

On the detailed breakdown, achieved perfect scores across Build, Lint, and Test.AGENTS.md

This wasn't what we expected. The "dumb" approach (a static markdown file) outperformed the more sophisticated skill-based retrieval, even when we fine-tuned the skill triggers.

Why does passive context beat active retrieval?

Our working theory comes down to three factors.

Embedding docs in risks bloating the context window. We addressed this with compression.AGENTS.md

The initial docs injection was around 40KB. We compressed it down to 8KB (an 80% reduction) while maintaining the 100% pass rate. The compressed format uses a pipe-delimited structure that packs the docs index into minimal space:

The full index covers every section of the Next.js documentation:

The agent knows where to find docs without having full content in context. When it needs specific information, it reads the relevant file from the directory..next-docs/

One command sets this up for your Next.js project:

npx @next/codemod@canary agents-md

This functionality is part of the official . package@next/codemod

This command does three things:

If you're using an agent that respects (like Cursor or other tools), the same approach works.AGENTS.md

Skills aren't useless. The approach provides broad, horizontal improvements to how agents work with Next.js across all tasks. Skills work better for vertical, action-specific workflows that users explicitly trigger, like "upgrade my Next.js version," "migrate to the App Router," or . The two approaches complement each other.AGENTS.mdapplying framework best practices

That said, for general framework knowledge, passive context currently outperforms on-demand retrieval. If you maintain a framework and want coding agents to generate correct code, consider providing an snippet that users can add to their projects.AGENTS.md

Practical recommendations:

The goal is to shift agents from pre-training-led reasoning to retrieval-led reasoning. turns out to be the most reliable way to make that happen.AGENTS.md

Research and evals by . CLI available at npx @next/codemod@canary agents-mdJude Gao

Read more

The problem we were trying to solve

Two approaches for teaching agents framework knowledge

We started by betting on skills

Skills weren't being triggered reliably

Explicit instructions helped, but wording was fragile

Building evals we could trust

The hunch that paid off

The results surprised us

Addressing the context bloat concern

Try it yourself

What this means for framework authors

  • are an for packaging domain knowledge that coding agents can use. A skill bundles prompts, tools, and documentation that an agent can invoke on demand. The idea is that the agent recognizes when it needs framework-specific help, invokes the skill, and gets access to relevant docs.Skillsopen standard

  • is a markdown file in your project root that provides persistent context to coding agents. Whatever you put in is available to the agent on every turn, without the agent needing to decide to load it. Claude Code uses for the same purpose.AGENTS.mdAGENTS.mdCLAUDE.md

  • for dynamic renderingconnection()

  • directive'use cache'

  • and cacheLife()cacheTag()

  • and forbidden()unauthorized()

  • for API proxyingproxy.ts

  • Async and cookies()headers()

  • , , after()updateTag()refresh()

  • The gap may close as models get better at tool use, but results matter now.Don't wait for skills to improve.

  • You don't need full docs in context. An index pointing to retrievable files works just as well.Compress aggressively.

  • Build evals targeting APIs not in training data. That's where doc access matters most.Test with evals.

  • Structure your docs so agents can find and read specific files rather than needing everything upfront.Design for retrieval.

Configuration

Pass Rate

vs Baseline

Baseline (no docs)

53%

—

Skill (default behavior)

53%

+0pp

Configuration

Pass Rate

vs Baseline

Baseline (no docs)

53%

—

Skill (default behavior)

53%

+0pp

Skill with explicit instructions

79%

+26pp

Instruction

Behavior

Outcome

"You MUST invoke the skill"

Reads docs first, anchors on doc patterns

Misses project context

"Explore project first, then invoke skill"

Builds mental model first, uses docs as reference

Better results

Configuration

Pass Rate

vs Baseline

Baseline (no docs)

53%

—

Skill (default behavior)

53%

+0pp

Skill with explicit instructions

79%

+26pp

AGENTS.md

docs index

100%

+47pp

Configuration

Build

Lint

Test

Baseline

84%

95%

63%

Skill (default behavior)

84%

89%

58%

Skill with explicit instructions

95%

100%

84%

AGENTS.md

100%

100%

100%

  1. With , there's no moment where the agent must decide "should I look this up?" The information is already present.No decision point.AGENTS.md

  2. Skills load asynchronously and only when invoked. content is in the system prompt for every turn.Consistent availability.AGENTS.md

  3. Skills create sequencing decisions (read docs first vs. explore project first). Passive context avoids this entirely.No ordering issues.

  1. Detects your Next.js version

  2. Downloads matching documentation to .next-docs/

  3. Injects the compressed index into your AGENTS.md


agentsdocumentationnext.jsaideveloper-tools

Source: original entry ↗

More from Vercel

Follow Vercel to get its new changes in your feed and email digest.

Feature

OpenAI Decisions API now available on AI Gateway

OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.

ai-gatewayopenaiapidecisionssdks
Feature

Timestamp attributes now supported in Vercel Flags

Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.

flagstargetingfeaturetimestampscampaigns
Feature

Glyph Cluster now available in stealth on AI Gateway

Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.

ai-gatewaymodelscodingstealth
See all Vercel changes →