megachangelog
Improvement

Making Turborepo 96% faster with agents, sandboxes, and humans

Turborepo now computes task graphs 81-91% faster through AI-driven optimizations including parallelization, allocation elimination, and syscall reduction. Time to First Task improved 11x on large monorepos through a systematic process combining coding agents, sandboxes, and traditional performance engineering.

is now 81-91% faster to compute its task graph in our repositories, scaling with repo size. On our 1,000+ package monorepo, now feels instant. Time to First Task is now 11x faster.Turborepoturbo run

After testing my changes with some open source Turborepos and asking Vercel customers to try canary releases on their repositories, I found the performance improvement could get as high as 96% depending on the size and complexity of the repository.

The process behind earning these performance gains is worth sharing, because it wasn't one optimization or one technique. It was eight days of mixing AI agents, Vercel Sandboxes, and typical, boring engineering practices.

Every starts by analyzing your monorepo's structure, scripts, and dependencies to build a task graph. That graph determines execution order, creates parallelism, and powers caching so you never repeat the same work twice.turbo run

Building the task graph is overhead you pay before your repository's work begins. The larger the repo, the higher the cost. On our 1,000-package monorepo, that cost was around 10 seconds on an M4 Pro Max. I don't know about you, but I found that unacceptable.

I wanted to see what agents could do about this without much guidance. I spun up 8 background coding agents from my phone before bed, each targeting a different part of the Rust codebase I suspected was too slow.

In each prompt, I replaced the part of the codebase I was interested in with a new target. I was curious what the agents would accomplish with plenty of ambiguity, as a baseline.

By morning, 3 of the 8 had produced outputs that I could turn into shippable wins:

These are undoubtedly meaningful successes, but reviewing all 8 chat sessions and code outputs taught me just as much about where unattended, state-of-the-art agents without proper context engineering will fall short today.

The agents running unattended produced some good wins, but I could tell this wouldn't be sustainable. We needed stronger testing, and a better verification loop. I had to be more involved.

The first normal engineering thing I did was take a profile. Shocking, I know.

I ran on our largest repo and opened the trace in .turbo run build --profilePerfetto

Flame graphs are informative, but can be slow to work with. As much as I do enjoy reading flame graphs and grinding out a win, Turborepo has a lot of shipping to do. I have a duty to users of Turborepo to work efficiently and effectively, using the best tools that I have at my disposal.

Turborepo's profiles are JSON files in Chrome Trace Event Format.

An LLM can theoretically read through and parse all this, but...well...just look at it. Function identifiers split across lines, irrelevant metadata mixed in with timing data, not grep-friendly. I pointed an agent at the file and watched it struggle through calls, trying to piece together function names from different lines, unsuccessfully trying to filter out noise. It was fumbling through this file in the same way I would.grep

One of my favorite heuristics for working with coding agents is that if something is poorly designed for me to work with, it's poorly designed for an agent, too. This isn't necessarily a comment about work quantity, but more so about interfaces. If something is hard for me to read, it stands to reason it's hard for an agent to read, too. This idea has its limits, but you'll see it quickly pay dividends in a moment.

A week prior, I saw about how Bun shipped a new flag: . It outputs profiles as Markdown, which easily fits into my view of how agents work best.a tweet from Jarred Sumner--cpu-prof-md

In , I added a new crate that generates a companion file alongside every trace. Hot functions sorted by self-time, call trees sorted by total-time, caller/callee relationships. All greppable, all on single lines.#11880turborepo-profile-md.md

The difference in the agent's output quality was dramatic. Same model, same codebase, same data, same agent harness. Different format, better optimization suggestions. The profile data was finally in a format that both I and the agent could read at a glance.radically

With Markdown profiles, I settled into a rhythm.

This loop produced over 20 performance PRs in four days. The wins fell into three categories. I'll give some examples.

was the largest. Building the git index, walking the filesystem for glob matches, parsing lockfiles, and loading files were all sequential operations that could run concurrently. PRs , , , and parallelized these hot paths.Parallelizationpackage.json#11889#11902#11927#11918

removed redundant copies and clones throughout the pipeline, including reference-based hashing in SCM operations (), pre-compiling glob exclusion filters (), and using a shared HTTP client instead of constructing a new one per request ().Allocation elimination#11916#11891#11929

batched per-package git subprocess calls into a single repo-wide index (), replaced git subprocesses with library calls (), and then replaced with the faster altogether ().Syscall reduction#11887#11938#11950libgit2libgit2gix-index

Again, it's typical, normal, boring software engineering stuff. I did try to turn this into a but it repeatedly made too many mistakes. The combination of the model, the harness, and the loop simply weren't dependable enough, and could move so much code out from underneath me too quickly. Maybe if I were working on a sideproject, I would have accepted it, but Turborepo powers some of the largest repositories in the world. I have to be fast responsible.Ralph Wiggum loopand

The most interesting pattern I noticed during this phase was how the codebase itself served as the agent's strongest feedback mechanism.

I'd point out a performance issue in code the agent was working on. We'd fix it together. Then I'd ask, "Do you see anywhere else where we can improve in the same way?" The agent would find more instances of the same pattern across the codebase. Depending on the size of the changes, I would either add the change to the PR or write it down to do later.

In places where the existing code had a sloppy pattern, the agent would write new code in the same style. Once I corrected one instance, the agent followed the correction going forward. In future conversations, without any memory or context carrying across chats, the agent would see the merged improvements in the source and stop reproducing the old patterns.

Over time, I noticed the agent spontaneously writing tests when I wasn't expecting it to. I saw it creating abstractions that matched what I would have done, which wasn't happening before. I would revisit a place in the codebase where the agent had previously been ineffective, and, with no changes to model or harness, it would produce better code outputs.

It turns out your own source code is the best reinforcement learning out there.

By the end of the week, Turborepo was roughly 85% faster on our largest repo. Before I started, I had arbitrarily set a goal of 95% better. The remaining gains were feeling within reach.

The problem became measurement. I had been running all benchmarks on my MacBook, and the reports were getting increasingly noisy. As the code gets faster, system noise matters more. Syscalls, memory, and disk I/O all have their variance.hyperfine

The profiles were noisy too. I had gotten the codebase to a point where the individual functions were fast enough that background activity on my laptop was drowning out any good signal.

Was the change I made 2% faster, or did I just get lucky with a quiet run? I couldn't confidently distinguish real improvements from noise. I needed a quieter lab for my science.really

are ephemeral Linux containers that only have what you put in them. No background daemons, no Slack notifications pulling CPU, no background programs making network requests. The machine's resources are entirely focused on what you're running.Vercel Sandboxes

I wrote a bash script that automated the entire benchmarking workflow. I'll put an abbreviated version of below.the full gist

You'll notice that, at the end of this script, I'm downloading the profiles back to my laptop. My agent could then inspect the benchmark results and Markdown profiles locally, and I could confidently tell whether a change was a real improvement or noise.

With clean signal from Sandbox, I could see real breakthroughs in low-level changes that were invisible on my noisy laptop.

Stack-allocated git OIDs ()#11984

Every file in the git index stored its 40-character SHA-1 hash as a heap-allocated . On our largest repo, alone was creating over 10,000 individual 40-byte heap allocations.Stringnew_from_gix_index

implements so existing consumers work unchanged, and means cloning is a 40-byte on the stack instead of a heap allocation. Profile data showed self-time dropped 15% and dropped 17%.OidHashDeref<Target=str>Copymemcpynew_from_gix_indexget_package_file_hashes_from_index

The most notable improvement across all three sizes was the reduction in run-to-run variance, which agrees with our theory of less allocator pressure and more predictable performance.

Syscall elimination ()#11985

Every cache fetch was performing three syscalls: , which returned , then , then . Weird pattern.stat(.tar)ENOENTstat(.tar.zst)open(.tar.zst)

After some digging, I figured out that the fallback existed for cache artifacts from Turborepo's Golang era (2021-2022). No modern version writes uncompressed cache entries, and cache entries rotate out constantly..tar

Across 962 cache fetches on our largest repo, self-time dropped from 200.5ms to 129.6ms, a 35% reduction.fetch

Move instead of clone ()#11986

The visitor dispatch loop was deep-cloning a from a precomputed map for each of roughly 1,700 tasks. Since each task ID appears exactly once in the dispatch stream, can move the value out at zero cost instead of cloning.(String, HashMap<String, String>)HashMap::remove()

After eight days, Time to First Task on our largest repo dropped from 8.1 seconds to 716 milliseconds.

I estimate this would have taken at least two months without agents, but I hope this article shows you that they didn't do the work for me. I was leading the entire time, deciding what to profile, which proposals to pursue, when to change tools, and when to change strategy. But the combination of my existing engineering knowledge, giving agents better tooling, and a clean benchmarking environment let me move at a pace that wouldn't have been possible six months ago.

These performance gains are now stable and ready for you to use. to learn more about the latest in Turborepo.Visit the Turborepo 2.9 release post

Read more

How Turborepo schedules your tasks

Starting with unattended agents

Making profiling work for agents and humans

Vercel Sandbox for benchmarking

Results

Released in Turborepo 2.9

  • netted a ~25% reduction in wall-clock time, reducing allocation pressure through hashing by reference instead of cloning an entire .PR #11872HashMap

  • replaced , one of our Rust dependency crates, with . A near 1:1 replacement that uses a faster hashing algorithm, creating a ~6% win.PR #11874twox-hashxxhash-rust

  • came from an existing comment that we hadn't gotten to yet. We needed to replace an unnecessary with a multi-source depth-first search (DFS). This wasn't on the hot path of , but my prompts didn't specify , did they? Fair.PR #11878Floyd-Warshall algorithmTODOturbo runwhich hot path

  • The agent never realized it could benchmark the improvements on the Turborepo codebase itself. Turborepo dogfoods Turborepo, so it could have easily built a binary and run it right on the source code to get end-to-end results.

  • The agent would hyperfixate on the first idea that it came up with and force it to work, rather than backing up and thinking abstractly about the problem (even though the chat logs showed it trying to do so).

  • The agent would chase the biggest number it could get, creating microbenchmarks that were relatively meaningless when it came to real-world performance. It would then crank out a 97% improvement for the benchmark, which actually amounted to a 0.02% real-world improvement.

  • Never once did an agent write a regression test.

  • Never once did an agent use the flag in the CLI.--profileturbo

Maybe Chrome Tracing JSON isn't the best format

Building LLM-friendly profiles

The iterative loop

Your source code is the best feedback loop

Hitting a wall at 85%

Breaking through the wall

  1. Put the agent in Plan Mode with instructions to create a profile and find hotspots in the Markdown output

  2. Review the proposed optimizations and decide which ones were worth pursuing

  3. Have the agent implement the good proposal(s)

  4. Validate with end-to-end benchmarkshyperfine

  5. Make a PR

  6. Repeat

Repo size

Before

After

Change

~1,000 packages

1.463s ± 0.052s

1.466s ± 0.027s

Same speed, 48% less variance

~125 packages

658.6ms ± 144.6ms

592.1ms ± 62.9ms

10% faster, 57% less variance

6 packages

96.8ms ± 46.7ms

75.0ms ± 18.4ms

22% faster, 61% less variance

Repo size

v2.8.0

v2.9.0

Improvement

~1,000 packages

8.1s

0.716s

91% faster

132 packages

1.9s

0.361s

81% faster

6 packages

0.676s

0.132s

80% faster

performanceturborepooptimizationmonorepoai

Source: original entry ↗

More from Vercel

Follow Vercel to get its new changes in your feed and email digest.

Feature

OpenAI Decisions API now available on AI Gateway

OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.

ai-gatewayopenaiapidecisionssdks
Feature

Timestamp attributes now supported in Vercel Flags

Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.

flagstargetingfeaturetimestampscampaigns
Feature

Glyph Cluster now available in stealth on AI Gateway

Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.

ai-gatewaymodelscodingstealth
See all Vercel changes →