SERHANT.'s playbook for rapid AI iteration with Vercel
SERHANT. scaled their AI-powered real estate platform from 200 pilot users to 900+ agents by using Vercel's AI SDK and AI Gateway to orchestrate multiple AI models (OpenAI, Claude, Gemini) for different tasks while maintaining infrastructure flexibility and reducing operational overhead. This approach enabled faster model switching, cost optimization, and seamless scaling without backend rewrites.
Impact at a glance
Using multiple models to balance cost, speed, and complexity
Worry-free scale: Adding users and assets
Future-proofing an unpredictable landscape
Started with Next.js on Vercel, which made it easier to expand to a React Native iOS app without rebuilding their backend
Engineers focus on AI design and iteration instead of platform plumbing
Orchestrates OpenAI, Claude, and Gemini by task to optimize cost vs output
Scaled from an internal pilot to 800–900+ real estate agents without replatforming
for complex, accuracy-critical analysis like comparative market analysis, where strong structured-data reasoning mattersClaude Sonnet
for lightweight intent and field-filling tasks where speed mattersClaude Haiku
for conversational voice and general chat behaviorsOpenAI models
for image generation, browser automation, and computer-use workflows where reliability and speed are the priorityGemini
When Jeremy Bunting joined SERHANT. as VP of Engineering in February 2024, was already showing promise. 200 real estate agents were piloting the AI product, which was designed to save time by automating cumbersome and repetitive daily tasks, like market analysis and contact management.S.MPLE
S.MPLE was a Next.js progressive web app deployed on Vercel, and that foundation gave the team leverage. They could keep the API layer steady while expanding the client experience, including expansion to a React Native iOS app, all without a backend rebuild.
But Bunting had a problem that keeps many engineering leaders up at night: the AI landscape changes faster than most teams can implement infrastructure updates.
The team needed to move fast, scale confidently, and stay flexible enough to swap models, add new capabilities, and adapt to the rapidly changing AI landscape. Traditional approaches meant choosing between velocity and flexibility, but Bunting wanted both.
As S.MPLE shifted from "one model” experiments to a production AI product, Bunting's team started evaluating , and he initially had concerns. "I asked, how much is this going to tie us in directly to Vercel?" he recalls.Vercel's AI SDK
Then one of his engineers pushed back. The AI SDK wasn't infrastructure lock-in, it was infrastructure independence. “It's just an SDK that abstracts away the complexity of working with different model providers”, the engineer pointed out.
Bunting also realized that if the team picked one frontier model and built tightly around it, every future change would come with a rewrite, with no clean path to fallback when reliability or cost shifted. With AI SDK, iteration meant simple configuration changes, not feature overhauls. "We are building agentic tools," said Bunting. "Having that consistent abstraction layer... really reduces the cognitive load."
added another layer of leverage: consolidated visibility into usage across apps and prototypes, even when teams bring their own keys. The result is faster debugging, faster optimization, and a clearer feedback loop on cost.AI Gateway
Because the SERHANT. S.MPLE team was not spending its time rebuilding infrastructure or maintaining one-off AI integrations, they could shift their attention to testing models against real product tasks and choosing the right tool for each job:
They are also experimenting with “models as guardrails” to validate or critique outputs, and with caching strategies to rein in token spend as usage grows.
The value of their stack decisions became clear when S.MPLE launched publicly. "We moved from being an internal pilot program to more than 900 users without a lot of worry on infrastructure or scale," Bunting says. The API layer didn't require a single change, and handled the increase in workloads automatically. Fluid compute
That seamless scale matters because SERHANT. operates at a content generation pace that would break most systems. "SERHANT. generates about 35% more content than the top five brokerages combined," Bunting notes. Between property videos, listing descriptions, marketing materials, and now AI-generated assets, the volume is staggering.
Greg Parsons, Technical Director on the S.MPLE team, said that AI Gateway gives him visibility across their platform that wasn’t possible before. "We can gain insight into all of the disparate applications we are building across the business," Parsons explained.
S.MPLE began with linear workflows: real estate agents would trigger a single action, it would run end-to-end then return the result.
But real-world workflows are more complex and users want to execute multiple tasks in a single run, like producing all of the assets needed to market a property listing. Bunting’s team is now building toward conversational experiences where humans can run agents, steer and correct mid-flight, and combine multiple “recipes” into a single request.
It is a shift from one-off automations to a coordinated network of specialized agents, built to evolve as fast as the AI ecosystem itself.
Greg Chan, SERHANT.’s CTO, sees flexibility as the point. "In AI, things are evolving fast. What it looks like now is different than even three-to-six months ago. And it’ll be different months from now," Chan says.
For Chan, the win is that the team can keep building inside the ecosystem instead of rewriting the stack every time the world changes. "The last thing we want is to rebuild our stack every time a new model drops," he said.
About SERHANT.
SERHANT. is an AI-native real estate and media company, and is the most-followed real estate brand in the world.
Founded in 2020 by top real estate broker and entrepreneur Ryan Serhant, SERHANT. brings together brokerage, media, and education, with proprietary technology to revolutionize how properties are marketed, sold, and experienced.
SERHANT. sells residential, commercial, luxury, and new development properties nationally through its specialized divisions, including SERHANT. Signature for high-net-worth clients and SERHANT. New Development, which delivers end-to-end branding, marketing, and sales for ground-up residential projects.
Powered by S.MPLE, SERHANT.’s proprietary AI platform, agents are empowered with real-time data and workflow automation across listings, transactions, and marketing to deliver faster, smarter, and more impactful results while saving time. Award-winning SERHANT. Studios produces original content across social and streaming platforms, while SellIt.com is the company’s digital education hub, engaging members globally in more than 130 countries.
Learn more at .serhant.com
AI SDK: Moving fast without vendor lock-in
What’s next: from linear workflows to conversational AI agents
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.