megachangelog
Feature

Run Harbor evals and Terminal-Bench on Vercel Sandbox

You can now run Harbor evaluations, including Terminal-Bench and other benchmarks like SWE-bench and OSWorld, on Vercel Sandbox. Each trial runs in an isolated Firecracker microVM with network policy enforcement and optional credential injection, allowing you to parallelize benchmarking beyond local machine capabilities and test multiple models via AI Gateway.

You can now run Harbor evals on Vercel Sandbox.

is the open-source harness behind , whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass to and each trial executes in its own isolated Firecracker microVM, so you can parallelize far beyond what your local machine is capable of. HarborTerminal-Bench--env vercelharbor run

A task's network policy is enforced at the sandbox firewall, outside the VM. Optional credential injection attaches secrets to matching outbound requests at that firewall, so they never enter the sandbox.

Paired with , one reaches hundreds of models from multiple providers, and benchmarking another model is the same command with a different :AI GatewayAI_GATEWAY_API_KEY--model

Swap to to run the same benchmark against an OpenAI model. --modelvercel_ai_gateway/openai/gpt-5.6-luna

Requires Harbor or later. Follow the for setup, configuration, and troubleshooting. Learn more in the .0.22.0step-by-step guideSandbox documentation

Read more

sandboxaibenchmarksharborevaluation

Source: original entry ↗