Evaluating bash vs SQL for AI agent querying
An evaluation comparing bash, SQL, and filesystem-based approaches for AI agents querying structured data found that SQL achieved 100% accuracy while bash achieved only 53%, though a hybrid approach combining both methods proved most reliable through self-verification. The open-source eval harness reveals that agent tool choice depends on the task: SQL for direct queries, bash for exploration and verification.
We invited from to share how they tested the "bash is all you need" hypothesis for AI agents.Ankur GoyalBraintrust
There's a growing conviction in the AI community that filesystems and bash are the optimal abstraction for AI agents. The logic makes sense: LLMs have been extensively trained on code, terminals, and file navigation, so you should be able to give your agent a shell and let it work.
Even non-coding agents may benefit from this approach. Vercel's recent post on showed this by mapping sales calls, support tickets, and other structured data onto the filesystem. The agent greps for relevant sections, pulls what it needs, and builds context on demand.building agents with filesystems and bash
But there's an alternative view worth testing. Filesystems may be the right abstraction for exploring and retrieving context, but what about querying structured data? We to find out.built an eval harness
We tasked agents with querying a dataset of GitHub issues and pull requests. This type of semi-structured data mirrors real-world use cases like customer support tickets or sales call transcripts.
Question complexity ranged from:
Three agent approaches competed:
Each agent received the same questions and was scored on accuracy.
SQL dominated. It hit 100% accuracy while bash achieved just 53%. Bash also used 7x more tokens and cost 6.5x more, while taking 9x longer to run. Even basic filesystem tools (search, read) outperformed full bash access, hitting 63% accuracy.
You can explore the , , and results directly.SQL experimentbash experimentfilesystem experiment
One surprising finding was that the bash agent generated , chaining , , , , and in ways that rarely appear in typical agent workflows. The model clearly has deep knowledge of shell scripting, but that knowledge didn't translate to better task performance.highly sophisticated shell commandsfindgrepjqawkxargs
The eval revealed substantive issues requiring attention.
Commands that should run in milliseconds were timing out at 10 seconds. calls across 68,000 files were the culprit. The addressing this.Performance bottlenecks.stat() tool received optimizationsjust-bash
The bash agent didn't know the structure of the JSON files it was querying. Adding schema information and example commands to the system prompt helped, but not enough to close the gap.Missing schema context.
Hand-checking failed cases revealed several questions where the "expected" answer was actually wrong, or where the agent found additional valid results that the scorer penalized. Five questions received corrections addressing ambiguities or dataset mismatches.Eval scoring issues.
The Vercel team submitted a .PR with the corrections
After fixes to both and the eval itself, the performance gap narrowed considerably.just-bash
Then we tried a different idea. Instead of choosing one abstraction, give the agent both:
The hybrid agent developed an interesting behavior. It would run SQL queries, then verify results by grepping through the filesystem. This double-checking is why the hybrid approach consistently hits 100% accuracy, while pure SQL occasionally gets things wrong.
You can explore the directly.hybrid experiment results
The tradeoff is cost. The hybrid approach uses roughly two times as many tokens as pure SQL, since it reasons about tool choice and verifies its work.
After all the fixes to , the eval dataset, and data loading issues, bash-sqlite emerged as the most reliable approach. The "winner" wasn't raw accuracy on a single run, but consistent accuracy through self-verification.just-bash
Over 200 messages and hundreds of traces later, we had:
The bash agent's tendency to check its own work turned out to be valuable just not for accuracy, but also for surfacing problems that would have gone unnoticed with a pure SQL approach.
For structured data with clear schemas, SQL remains the most direct path. It's fast, well-understood, and uses fewer tokens.
For exploration and verification, bash provides flexibility that SQL can't match. Agents can inspect files, spot-check results, and catch edge cases through filesystem access.
But the bigger lesson is about evals themselves. The back-and-forth between Braintrust and the Vercel team, with detailed traces at every step, is what actually improved the tools and the benchmark. Without that visibility, we'd still be debating which abstraction "won" based on flawed data.
The .eval harness is open source
You can swap in your own:
This post was written by and the team at , who build evaluation infrastructure for AI applications. The eval harness is open source and integrates with from Vercel.Ankur GoyalBraintrustjust-bash
Setting up the eval
Initial results
Debugging the results
The hybrid approach
Key learnings
What this means for agent design
Run your own benchmarks
Simple queries: "How many open issues mention 'security'?"
Complex queries: "Find issues where someone reported a bug and later someone submitted a pull request claiming to fix it"
"Which repositories have the most unique issue reporters" was ambiguous between org-level and repo-level grouping
Several questions had expected outputs that didn't match the actual dataset
The bash agent sometimes found more valid results than the reference answers included
Let it use bash to explore and manipulate files
Also provide access to a SQLite database when that's the right tool
Fixed performance bottlenecks in
just-bashCorrected five ambiguous or wrong expected answers in the eval
Found a data loading bug that caused off-by-one errors
Watched agents develop sophisticated verification strategies
Dataset (customer tickets, sales calls, logs, whatever you're working with)
Agent implementations
Questions that matter to your use case
: Direct database queries against a SQLite database containing the same dataSQL agent
: Using to navigate and query JSON files on the filesystemBash agent
just-bash: Basic file tools (search, read) without full shell accessFilesystem agent
Agent | Accuracy | Avg Tokens | Cost | Duration |
|---|---|---|---|---|
SQL | 100% | 155,531 | $0.51 | 45s |
Bash | 52.7% | 1,062,031 | $3.34 | 401s |
Filesystem | 63.0% | 1,275,871 | $3.89 | 126s |
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.