35%
“35% lower cost per million tokens”
As published on baseten.co, July 2024. Captured by usedby on Oct 5, 2026.
What happened
Writer, a generative AI platform for enterprises, uses Baseten to serve its custom 70-billion-parameter Palmyra language models in production. Baseten's performance engineers built optimized TensorRT-LLM engines for secure, private deployments of the healthcare and finance models.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Writer, not ours.
- 60%“60% higher tokens per second”
- 23%“23% lower time to first token”
With TensorRT-LLM engines deployed on Baseten, Writer surpassed its performance requirements ahead of launching the new models.
Inference for custom-built LLMs could be a major headache. Thanks to Baseten, we’re getting cost-effective high-performance model serving without any extra burden on our internal engineering teams. Instead, we get to focus our expertise on creating the best possible domain-specific LLMs for our customers.
What the story claims, and what we checked
We compared the story with its live page on Oct 5, 2026.
- The figure: 35%CheckedPrinted word for word on the page, near the name of Writer.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Writer.
- The numbers in our summaryCheckedEach one is printed on the page.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Writer uses BasetenCheckedConfirmed line. Latest check across sources: Oct 5, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




