The Daily Model Drift Vol. 0 · Pre-launch In calibration ·

A daily public record of how AI models change

Models change quietly. We’re about to write it down.

Daily Model Drift will test AI models every day through the tools people use them in, grade every answer with code, and publish every response. When a model quietly changes, there will be a dated record.

  1. GPT-6 Luna
  2. GPT-6.1 Sol
  3. Claude Sonnet 5.5
  4. Claude Haiku 4.5
  5. Qwen 3.8 27B
  6. Gemini 3.8 Flash · adapter pending
Fig. 1. Calibration trace. Each line is one model’s channel: a reference pulse, then the pen settles to zero. It is drawn, not measured. The first real readings appear at launch.

Instrument status

Baseline collection starts soon.

There is no launch date yet. Once daily collection begins, the first verdicts need about 20 to 25 days of history, and the tests reach full strength at about 35.

Pilot of real calls
Passed, 4 October 2026. One live call per model, graded by code. Gemini’s adapter blocked its call.
Test suite, version 2
Calibration under way: tuning each model’s task difficulty.
Daily baseline
Not started yet.
Gemini 3.8 Flash
Pending a decision on its adapter.
Launch date
Not set.
Published measurements
None yet. Nothing is shown until it is real.

I. Why this exists

The model behind your tool is not a fixed thing.

Vendors adjust serving systems, safety layers, default effort and the tools themselves, usually without a new version number. People notice: it got lazier, it started refusing, it stopped finishing files. There is rarely a record to check that against.

In September 2025 Anthropic published a postmortem of three serving bugs. At its peak one of them affected 16% of Claude Sonnet 4 requests, and the company’s own evaluations had not caught it. Changes like this are real, intermittent and hard to see from inside one conversation.

Daily Model Drift will keep that record: the same kinds of task, fresh every day, graded the same way, published in full.

II. Method

Small tests, every morning, graded by code.

  1. Fresh tasks, every day

    Generated from a seed in nine families, so no prompt repeats. Difficulty is tuned per model to pass about 60–75% of the time, leaving room for a drop to show.

  2. Through the real tools

    Each model runs through its vendor’s own command-line tool, the way people use it: Claude Code, Codex CLI, Antigravity. Qwen goes through OpenRouter.

  3. Graded by code

    Exact integers, line-by-line diffs, counted constraints. No model grades another, and no model-written code is run.

  4. Compared with itself

    Each day is tested against the model’s own previous 28 days, and each week against the four weeks before it.

  5. A verdict, with evidence

    Collecting baseline, stable, watch or drift, each linked to the prompts, raw responses and grades behind it.

The nine families arithmetic chains · code tracing · key–value lookup · constraint following · JSON extraction · whole-file edits · long enumerations · benign requests with alarming words · natural answer length.

III. Specimens

Three items, in the form the generators write them.

Illustrative. Real items are generated fresh each day and published after the run.

arith-chain

Compute the value of ((38 * 7) - 145) * 3 + 12. Reply with only the integer.
Graded by
Exact match against 375, computed by our code. One integer, nothing else.
Watches for
Less reasoning effort: chains that used to be worked through get guessed.

full-file-edit

Here is src/ledger.js (118 lines). Make the following changes:
- Rename parseRow to parseRecord.
- Change MAX_ROWS from 500 to 750.
Output the entire modified file, every line.
Graded by
Line-by-line diff against the expected file. Placeholders such as “// rest unchanged” are counted.
Watches for
Lazy edits: elided code, placeholder comments and files cut short.

long-enumeration

Define f(i) = 3i + 7. List f(i) for every i from 41 through 160 (120 values). Format each line as `i: f(i)`. Output every line in full.
Graded by
Every line checked. We record how far it got and whether it stopped early or wrote “…”.
Watches for
Truncation: answers that stop before the job is done.

IV. House rules

What we will never do with the numbers.

  1. Real measurements only.

    This page shows none yet, and the site will never present invented or synthetic numbers as measurements. The release build refuses to run while any synthetic data exists.

  2. Unknown stays unknown.

    A timeout, a harness error or a missing usage figure is recorded as an error with its reason, never as a pass or a fail. Gaps are not filled in afterwards.

  3. Every response is evidence.

    Each prompt, raw response, grade and tool version will be published at a dated permalink, so any verdict can be checked by hand.

  4. A model is compared only with itself.

    Difficulty is tuned per model, so scores are not comparable across models and there is no leaderboard.

  5. We say what it cannot see.

    Every verdict comes with the size of change the test could have missed.

V. Instruments

Six models, each through its own tool.

Models planned for daily testing
ModelCalled throughStatus
GPT-6 Luna low effortCodex CLIPlanned
GPT-6.1 Sol high effortCodex CLIPlanned
Claude Sonnet 5.5Claude CodePlanned
Claude Haiku 4.5Claude CodePlanned
Qwen 3.8 27B free tierOpenRouter, through Claude CodePlanned
Gemini 3.8 Flash low effortAntigravityPending an adapter decision

Effort settings are pinned, and each vendor tool’s version is recorded next to every result. A new tool version pauses that model’s tests for seven days while the baseline restarts.

VI. Follow the launch

One feed. No sign-up.

feed.xml Atom

Add https://dailymodeldrift.com/feed.xml to any feed reader. It has one entry now and will get a few more: when baseline collection starts, when a launch date is set, and at launch. After launch the same feed carries flagged drift events, and nothing else.

No newsletter, no email list, no cookies, no analytics. The data will be open: every result downloadable as JSON at a stable address. The data licence is still being decided.