As published on wandb.ai, November 2025. Captured by usedby on Oct 8, 2026.
What happened
Shell's NLP research team uses Weights & Biases to log experiments, build custom dashboards, run hyperparameter sweeps, and track evaluations and user feedback while adapting large language models to its internal technical reports. The resulting models power a research assistant that acts as an institutional memory for researchers.
Summary written by usedby from the source page, in English. The figures are those of Weights & Biases and Shell, not ours.
- 26%“The DAPT + instruction-tuning models led to a 26% increase in accuracy compared to non-fine-tuned Llama models, and a 30% improvement in domain-specific understanding on Shell benchmarks.”
At the core of this scalable infrastructure is Weights & Biases, which plays critical roles in supporting experiment tracking and logging, building custom dashboards for increased visibility into training dynamics, robust hyperparameter optimization, and evaluation and feedback through both W&B Models and W&B Weave.
We built a custom dashboard using Weights & Biases for all our experiment logging, and a lot of critical hyperparameter tuning using W&B Sweeps to explore impact and sensitivity.
What the story claims, and what we checked
We compared the story with its live page on Oct 8, 2026.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Shell.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Shell uses Weights & BiasesCheckedConfirmed line. Latest check across sources: Oct 8, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




