<200ms
latency for inline edit suggestions in production
As published on baseten.co, March 2026. Checked by usedby on Oct 10, 2026.
What happened
Posit hosts and serves the small fine-tuned LLMs behind its Next Edit Suggestions feature in RStudio on Baseten's inference stack, and used Baseten Training to run fine-tuning experiments before launching Posit AI.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Posit, not ours.
By deploying and optimizing the model on the Baseten Inference Stack, Posit achieved sub-200ms latencies, which was the critical threshold for the inline edit suggestion experience.
Being able to stand up compute in a matter of minutes helped us to control costs while rapidly iterating on our training runs.
What the story claims, and what we checked
What we compared with the page.
- The figure: <200msCheckedThe figure is on the page; its label is our wording.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Posit.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Posit uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




