Skip to main content
usedby

<200ms

latency for inline edit suggestions in production

As published on baseten.co, March 2026. Checked by usedby on Oct 10, 2026.

What happened

Posit hosts and serves the small fine-tuned LLMs behind its Next Edit Suggestions feature in RStudio on Baseten's inference stack, and used Baseten Training to run fine-tuning experiments before launching Posit AI.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Posit, not ours.

By deploying and optimizing the model on the Baseten Inference Stack, Posit achieved sub-200ms latencies, which was the critical threshold for the inline edit suggestion experience.

From the page. baseten.co

Being able to stand up compute in a matter of minutes helped us to control costs while rapidly iterating on our training runs.

Simon Couch, AI Core Team, Posit. Source

What the story claims, and what we checked

What we compared with the page.

  • The figure: <200msCheckedThe figure is on the page; its label is our wording.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Posit.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Posit uses BasetenCheckedConfirmed line.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

30%“30% lower cost per call”you.com uses Baseten. Another customer of Baseten3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems uses Baseten. Another customer of Baseten33%“High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers.”OpenCode uses Baseten. Another customer of Baseten