If a dbt model or an Airflow task calls an LLM, treat the call like a vendor API pull. Land each response once, key it on sha256 of the input, the model ID, and the prompt version, and let backfills read the stored rows. Only a new key should ever call the model.
This is the entire post. Everything below is why it is not optional.
Temperature 0 was never deterministic
Regenerating will not match. That has been documented since at least 2024 (arXiv 2408.04667), and in September 2025 Thinking Machines Lab pinned down why: one prompt, 1,000 runs at temperature 0, 80 unique completions, because batch size shifts with server load. The nondeterminism is not in your prompt. It is in the serving infrastructure, and you do not control it.
So every rerun of an LLM step is a new sample, not a replay. A backfill that re-calls the model does not reproduce history. It rewrites it, one row at a time, with no record of what the old values were.
The model IDs are moving too
What is new in 2026: on the Claude API, Claude 4.7 and later models return a 400 for any temperature except the default 1.0. Retired model IDs stop answering entirely; claude-sonnet-4-20250514 was retired June 15, 2026. Anthropic lists tentative retirements with at least 60 days’ notice, but the direction is one way. The model you called last quarter may not answer next quarter.
Before an upgrade, search your code and configs for temperature: kwargs, dicts, YAML, all of it, and remove the parameter. The Anthropic Python SDK v1.0+ drops temperature entirely, so passing it raises a TypeError. The parameter is dead. Bury it.
The architecture
The model call lives outside dbt. Airflow, or your orchestrator, calls the API and lands the raw response. The dbt model is an incremental staging model that dedupes by the key. Backfills read stored rows. The key is sha256 of input, model ID, and prompt version, which means a model migration or a prompt change naturally produces new keys and new calls, while an ordinary backfill touches nothing.
This is the dbt reference card in one paragraph: an incremental config on the staging model, the llm_key spec above, and the backfill-twice test as the proof. If the second backfill calls the API or changes a row, the key is wrong.
Why this is a research-to-production story
The nondeterminism finding is research. The keying pattern is production engineering. The gap between them is where most teams live: they read the temperature-0 papers, nod, and keep rerunning LLM steps that silently rewrite history.
Three failure modes to check against your own pipelines: label drift from assuming temperature 0 is deterministic, a breaking API parameter deprecation like the temperature removal, and a retired model ID that stops answering mid-quarter. The keying pattern survives all three. The unkeyed pipeline survives none of them.
This post started as an Instagram post →
Numbers above trace to these sources. If one moved, tell us and we fix it.


Talk it through
Argue with us on Instagram.