RESEARCH TO PRODUCTION

Rerunning an LLM step changes the output. Land it once, key it, and stop paying twice.

Temperature 0 was never deterministic. Treat every LLM call like a vendor API pull: land each response once, keyed on sha256 of input, model, and prompt version, and let backfills read the stored rows.

2026-09-29

If a dbt model or an Airflow task calls an LLM, treat the call like a vendor API pull. Land each response once, key it on sha256 of the input, the model ID, and the prompt version, and let backfills read the stored rows. Only a new key should ever call the model.

This is the entire post. Everything below is why it is not optional.

Temperature 0 was never deterministic

80
unique completions from 1,000 runs of one prompt at temperature 0, measured by Thinking Machines Lab in September 2025. The cause: batch size shifting with server load.

Regenerating will not match. That has been documented since at least 2024 (arXiv 2408.04667), and in September 2025 Thinking Machines Lab pinned down why: one prompt, 1,000 runs at temperature 0, 80 unique completions, because batch size shifts with server load. The nondeterminism is not in your prompt. It is in the serving infrastructure, and you do not control it.

So every rerun of an LLM step is a new sample, not a replay. A backfill that re-calls the model does not reproduce history. It rewrites it, one row at a time, with no record of what the old values were.

The model IDs are moving too

What is new in 2026: on the Claude API, Claude 4.7 and later models return a 400 for any temperature except the default 1.0. Retired model IDs stop answering entirely; claude-sonnet-4-20250514 was retired June 15, 2026. Anthropic lists tentative retirements with at least 60 days’ notice, but the direction is one way. The model you called last quarter may not answer next quarter.

Before an upgrade, search your code and configs for temperature: kwargs, dicts, YAML, all of it, and remove the parameter. The Anthropic Python SDK v1.0+ drops temperature entirely, so passing it raises a TypeError. The parameter is dead. Bury it.

The architecture

The model call lives outside dbt. Airflow, or your orchestrator, calls the API and lands the raw response. The dbt model is an incremental staging model that dedupes by the key. Backfills read stored rows. The key is sha256 of input, model ID, and prompt version, which means a model migration or a prompt change naturally produces new keys and new calls, while an ordinary backfill touches nothing.

This is the dbt reference card in one paragraph: an incremental config on the staging model, the llm_key spec above, and the backfill-twice test as the proof. If the second backfill calls the API or changes a row, the key is wrong.

Why this is a research-to-production story

The nondeterminism finding is research. The keying pattern is production engineering. The gap between them is where most teams live: they read the temperature-0 papers, nod, and keep rerunning LLM steps that silently rewrite history.

Three failure modes to check against your own pipelines: label drift from assuming temperature 0 is deterministic, a breaking API parameter deprecation like the temperature removal, and a retired model ID that stops answering mid-quarter. The keying pattern survives all three. The unkeyed pipeline survives none of them.

This post started as an Instagram post →

Sources

Numbers above trace to these sources. If one moved, tell us and we fix it.

Discussion

Talk it through

Argue with us on Instagram.