# Daily Model Drift > Daily Model Drift will test AI models every day through the tools people use, grade every answer with code and publish every response as evidence. It is in calibration now; no measurements are published yet. Status as of 2026-10-04: pre-launch. This site publishes no measurements yet, and nothing on it should be cited as a result. - A pilot of one real call per model passed on 2026-10-04. The second version of the test suite is being calibrated. - Daily baseline collection: not started. Verdicts need about 20 to 25 days of history; full detection power at about 35 days. - Launch date: not set. ## What it will measure Every day each model gets fresh, seeded tasks in nine families (arithmetic chains, code tracing, key-value lookup, constraint following, JSON extraction, whole-file edits, long enumerations, benign requests with alarming words, natural answer length), called through the vendor's own command-line tool. Code grades every answer; no model grades another. Each model is compared only with its own previous 28 days, and each week with the four weeks before it. Verdicts: collecting baseline, stable, watch, drift. ## Models - GPT-6 Luna (low effort), through Codex CLI: planned - GPT-6.1 Sol (high effort), through Codex CLI: planned - Claude Sonnet 5.5, through Claude Code: planned - Claude Haiku 4.5, through Claude Code: planned - Qwen 3.8 27B (free tier), through OpenRouter, through Claude Code: planned - Gemini 3.8 Flash (low effort), through Antigravity: pending an adapter decision ## Rules - Real measurements only; no invented or synthetic numbers are published as measurements. - Unknown stays unknown: timeouts, harness errors and missing usage are recorded as errors, never as passes or fails. - Every prompt, raw response, grade and tool version will be published at a dated permalink. - The data will be open; the licence is not decided yet. ## Links - [Home](https://dailymodeldrift.com/) - [Atom feed](https://dailymodeldrift.com/feed.xml): launch notes, then flagged drift events