2–3×
“2–3× lower infrastructure costs”
As published on together.ai. Checked by usedby on Oct 9, 2026.
What happened
XY.AI fine-tunes Qwen 2.5 14B LoRA adapters per customer on Together AI's managed fine-tuning platform to parse healthcare Explanation of Benefits documents into structured JSON. The adapters are served on serverless inference, with token log probabilities used to route low-confidence predictions to human review.
Summary written by usedby from the source page, in English. The figures are those of Together AI and XY.AI Labs, not ours.
- 87%“87% EOB parsing accuracy”
- 77%“driving EOB‑to‑JSON parsing accuracy from 77% to 87% when measured against expert human baselines”
- 10–20 minutes“On Together’s managed infrastructure, training runs complete in 10–20 minutes at roughly $10 per run.”
- 3ד3× faster training cycles”
Together’s fully managed training and serving eliminated the need for a dedicated AI infrastructure hire. XY.AI’s existing team handles all model development, while infrastructure costs dropped 2–3× compared to the prior self‑hosted setup
Together AI does for fine-tuning and inference what Vercel does for LLM-based apps—it removes the infrastructure layer so we can focus on our product. We fine‑tune and deploy customer‑specific models through simple API calls. That lets our existing team move from weekly to daily iteration, cut costs by 2–3×, and improve accuracy from 77% to 87%.
What the story claims, and what we checked
What we compared with the page.
- The figure: 2–3×CheckedPrinted word for word on the page, near the name of XY.AI Labs.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of XY.AI Labs.
- The numbers in our summaryCheckedEach one is printed on the page.
- XY.AI Labs uses Together AICheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.







