Improvement0.0.86
SDK v0.0.86
This release improves context-window handling for local OpenAI-compatible servers, adds streaming transcription support across multiple providers, and introduces a new shared cloud session runtime in @cline/core. It also fixes session persistence issues, hub startup error reporting, and request cancellation timing, updates the plugin command service as a shared core, and refreshes the model catalog with 6,386 total models.
- A text-only turn truncated at the output-token limit now gets one compact-and-retry attempt per run before the existing nudge-and-retry recovery. Local OpenAI-compatible servers (llama.cpp, Ollama, LM Studio) cap generation at whatever context remains regardless of the requested output budget, so long sessions died mid-answer even with sensible limits. The runtime now routes that finish through the same forced-compaction path
prepareTurnuses for context-window overflow; if compaction has nothing to remove or the retry truncates again, the loop's nudge-and-retry takes over and the original max-tokens error still surfaces with the partial answer persisted. Turns that produced tool calls, or provider-executed tool activity, are never replayed. A newtask.max_tokens_recoverylifecycle event (started,retried,failed) measures how often it helps - Hub startup failures now say why.
ensureCompatibleLocalHubUrlswallowed the error fromensureDetachedHubServer, so every failure read as "No compatible hub runtime is available"; the cause is now appended to that message and attached ascause, and the daemon reports its own startup failure (hub.daemon.startup) before flushing telemetry. The wait for a freshly spawned hub also goes from 8 to 15 seconds for every client, because the first launch after an install or update regularly takes 8 to 13 seconds on Windows while the new binary is scanned - Session renames made through the hub now persist.
HubRuntimeHost.updateSessionfolded the new title intometadata.title, but the persistence layer only treats an explicittitleas a rename, so the name reverted on relaunch, and wholesale metadata replacement dropped other keys such as pinned state.session.updatenow carriespromptandtitleas their own fields and leaves untouched metadata alone - Errors shown in a session now survive navigation and resume. Terminal failures, including those recorded after retries are exhausted, are persisted as display-only history entries that hosts render when a session is reopened, and those entries are excluded from compaction and from what the model sees
- Plugin slash commands are now a shared core service.
createPluginCommandServicein@cline/coreowns plugin command discovery, loading, name and result normalization, and execution, and both the CLI and the desktop sidecar consume it. A plugin that fails to load (an invalid manifest, a sandbox startup timeout) no longer rejects the service: the failure is logged, the empty command set is cached for the current plugin set, and the load is retried after 30 seconds instead of spawning a sandbox on every prompt. Handler exceptions still propagate - Model lists for endpoint-owned providers now report failures instead of returning an empty list. For private-catalog providers (LiteLLM, Baseten, Hicap, Poolside) and providers backed by a
modelsSourceUrl(Ollama, LM Studio), an unreachable host, TLS rejection, or bad key now propagates fromgetLocalProviderModelswhen the caller asks for errors, so hosts can show the cause rather than "no models". The hub guards its initial model load so a failing default provider no longer aborts peer setup - Aborting a request now takes effect during the empty-response backoff. The retry middleware slept through its backoff without watching the
AbortSignaland then re-dialed with an already-aborted signal, so cancellation surfaced up to one backoff late.createGatewayApiHandlerAsyncalso ignoredsetAbortSignal()because it captured the signal at construction; it now reads the live one like the sync handler - The yolo-mode system prompt replaces the vague "always show your planning process" rule with concrete output guidance: keep plans to one short paragraph, act once the next step is clear, put code and edit payloads in tool arguments rather than drafting them in text, skip preambles for routine tool calls, and resolve repeated uncertainty with a focused check instead of more speculation
- New built-in provider
aiand(ai&), an OpenAI-compatible endpoint serving open-weight models from Japan. It readsAIAND_API_KEYand defaults tozai-org/glm-5.3 - Streaming transcription in
@cline/llmsnow covers native OpenAI (gpt-realtime-whisperwith a transcription-bound client secret), Vercel AI Gateway, and ElevenLabs, exposed through the newgetBuiltinStreamingTranscriptionModels. Gateway and OpenAI-compatible batch transcription use AI SDK transcription with retries, cancellation, andproviderOptions, detect the audio format from the bytes, and returnsegmentsandwarningsalongside the text. Voice model discovery validates audio-to-text capability and keeps realtime-only models separate @cline/corehas a new./cloudsubpath export with the shared cloud session runtime (API client, controller, snapshots, state) that hosts previously each implemented. The cron spec watcher also resolves its directory withrealpathSync.nativebefore watching, avoiding a fatal libuv assertion on Windows when the path contains a short-name segment likeRUNNER~1@cline/sharedexportstoPosixSeparators, replacing four copies of the same helper across core and the VS Code extension. No behavior change- Refreshed the model catalog: still 209 providers, 6,237 to 6,386 models. The bundled Cline catalog drops
z-ai/glm-5.3-flash,cline-free/solar-pro4, andpoolside/laguna-s-2.1:free, and addscline-free/mimo-v2.6-flash. The resolved default model changes for 19 providers that do not pin one inbuiltins.ts: 11 land on Claude Opus 5.5 (Cortecs, CrossModel, DigitalOcean, Eden AI, GitHub Copilot, both LLM Gateway providers, Ofox, Requesty, Vertex, Vivgrid), both StepFun providers move to Step 5 Preview, Above to MiMo V2.6 Flash, Fireworks to Ember-1, Kenari to DeepSeek V4.1 Flash, NanoGPT to Aion 3.5, OpenCode Go to Space Bunny Free, and Pioneer to GLiNER 2.5 Multi. If you use one of those providers without pinning a model, expect a different default
Full Changelog: sdk/sdk/v0.0.85...sdk/sdk/v0.0.86
sdkcontext-windowtranscriptionlocal-modelsproviderreliability
Source: original entry ↗