systemcards.org / analysis

Watching AI model documentation:
what changes, what goes quiet, and what the watcher misses

cardtrack is a daily tracker of first-party model and system cards, access policies and independent evaluations of frontier AI models, with additions proposed by a large language model (LLM) agent and audited by a human (systemcards.org). We analyse its first 7.3 weeks: 323 documents from 33 publishers, 594 stored versions and 17,203 link and content checks. Post-publication edits are real and worth tracking: 61 publisher-side version changes were substantive, among them 14 corrections of a stated number or fact (one propagated through 5 OpenAI documents), scope changes to access policies, and score edits and removals with no change log; 49 documents (15.2%) changed substantively within the window. The cost of finding them is that 58.0% of detected changes were page furniture or extraction artefacts, and apparent deletions must be checked against the raw capture before they are believed. Hard loss is rare: 2 of 346 documents ever returned a 404, one of them for good, too few to estimate a link-rot rate over this window. The quieter failure is a page that keeps answering HTTP 200: two pages served a redirect stub for at least 17 days, visible only to content fingerprinting until the publisher added real redirects. A host-side fault cost 41.5% of the daily runs their checks, none since 22 September (Appendix B), so the change counts above are lower bounds and their detection dates run late.

Setting and data

Each daily run probes every tracked URL (link check), re-fetches a 15 % rotation of documents and compares a text fingerprint that ignores known page furniture such as download counters, sidebars and footers (content check), diffs publisher index pages for new links, and then lets a sandboxed LLM agent propose additions through a deterministic validator. A supervised backfill on 9 August 2026 seeded 196 documents; the snapshot covers 53 runs from 11 August to 29 September 2026, during which 159 more were discovered. The only continuous public log of edits to such documents we know of is the Midas Project’s hand-curated Watchtower (The Midas Project 2026); this study is fingerprint-based and model-centred.

What changes

(a) Classification of stored version pairs by what the text diff contained (groups with ten or more pairs; the “before filter fix” row lacks the versions later deleted; see text). (b) Kaplan–Meier estimate, that is the share of documents with at least one substantive publisher-side change by day tt adjusted for documents tracked fewer than tt days, by the document’s current format; changes are dated at detection, and backfilled documents were already months old on day 0.

Of the 239 consecutive version pairs in the database, 32 (13.4%) are the tracker’s own doing, so the diff compares two captures, not two revisions: canonical URLs migrated on 31 August from announcement pages to full PDFs, and first captures of a redirect target after a recorded move. The remaining 207 publisher-side pairs were classified by reading the text diffs (Figure 1a): 58.0% were Hugging Face download counters, evaluation-widget rows, footers or PDF re-extraction artefacts, 29.5% changed the document’s content and 8.7% did so substantially. On 31 August the fingerprint was also recomputed to ignore known furniture. This cut the per-check change-flag rate from 49.2% to 17.6% while the number of substantive pairs found did not fall (23 before, 38 after, over a longer window), so precision per flag rose from about 13.1% to 24.1%. The before side is incomplete: 107 versions that the changelog records as written are no longer in the database, 106 of them from before the recompute, so pairs of every kind are missing there. The remaining leak is evaluation-widget rows that end in a score, which the bullet-anchored patterns do not match. In document terms, 49 of 323 tracked documents (15.2%) changed substantively during the window, an estimated 16.5% of HTML and 15.9% of PDF documents within 30 days of tracking; the two curves in Figure 1b are not distinguishable at this sample size.

The substantive changes are the reason to run a tracker at all (Table 1; Appendix A lists all 61 with URLs). Of the publisher-side pairs, 14 corrected a stated number or fact, 13 touched an evaluation number and 24 touched a safety, risk or mitigation section. Three patterns recur (Appendix A ids in brackets). First, some corrections come with dated change logs, mostly at OpenAI: one corrected number propagated with its change log into 5 OpenAI documents that quote it as a baseline (A4, A11, A14, A19, A22), and most recently the GPT-6 Astra system card gained an appendix covering GPT-6 Sol and Luna and corrected Astra’s HealthBench results after an evaluation misconfiguration, with dated entries (A52); xAI (A12) and Anthropic (A44) each added one dated log. Second, edits without a change log exist: an NVIDIA card revised a tool-calling score downward and dropped its research-only use restriction with no note (A6); OpenAI removed a prompt-response table in which GPT-5.6 Cyber, alone of four configurations, answered a request for a Keychain-bypass and Chrome-cookie-decryption tool (A27), and deleted a named launch partner from an access-programme page (A40). Third, documents change scope: OpenAI rebuilt its Daybreak cyber-access tiers to add GPT-6 Sol, Luna and Astra and to require extra approval for GPT-5.6-Cyber (A54); Anthropic’s cyber-safeguards page now excludes Opus 5.5 and Sonnet 5.5 (A59); NVIDIA’s Cosmos 3 cards were each cut back to a single model (A50, A51); and inclusionAI added a “training content summary” linking a public training-data statement to several of its cards (A53, A55, A56, A57, A58). A caution the other way: the previous edition found that apparent deletions of Anthropic footnotes were extractor artefacts. The publisher had wrapped its footnotes in an element whose class name contains “footer”, which the boilerplate remover discards, while the raw HTML still held every footnote. The extractor has since been fixed and twelve classified versions, the artefacts among them, were deleted from the database; one new footnote deletion in the same position is treated as an artefact until it is checked against the raw capture. A fingerprint tracker cannot tell disclosed from undisclosed edits, or edits from extraction bugs, but it produces the diff that lets a human tell.

Publisher-side changes selected by rule from Appendix A (edits the classifier described as undisclosed, dated corrections, pages whose content moved behind an HTTP 200, licence changes, then the largest major edit), one row per distinct change, capped per kind. Summaries are the classifier’s; URLs and version ids are in Appendix A.
Date Publisher: document Kind What changed (classifier summary)
2026-08-19 nvidia: NVIDIA-Nemotron-3.5-Lightning-30B-A3... no change log Added Hardware Matrix, deployment table, SGLang recipes and speculative-decoding flags. Dropped RTX 5090 from supported hardware and removed the runtime thinking-budget claim.
2026-08-22 xai: Model Card: Grok 4.6 dated correction Revision 2026-08-17 adds changelog, PartBench, DeepSearchQA, BixBench, KernelBenchInternal v1.1, renumbered sections; corrects HackerBench, self-harm, MASK, LAB scores.
2026-08-25 tencent hunyuan: EVIE-Preview-4.5B Model Card no change log Card rewritten around ’Rank #1’ claims: new ViDoRe V3 table with 1,792-token tier (65.36), revised V1+V2 and per-domain scores, Index Cost section.
2026-08-26 nvidia: NVIDIA Nemotron Parse 2.0 Model Card licence License changed from NVIDIA Open Model License Agreement to OpenMDW-1.1; third-party software notice added; Quick Start and vLLM instructions rewritten.
2026-09-06 palisade research: Language Models Can Autonomously Hac... moved behind 200 Blog URL now serves a meta-refresh stub ’Redirecting... Click here’ pointing to palisaderesearch.org/research/self-replication; article text no longer captured.
2026-09-06 palisade research: Technical Report: Shutdown Resistanc... moved behind 200 Page body (shutdown-resistance-on-a-robot report: 3/10 physical, 52/100 simulated trials) replaced by a two-line redirect stub: Redirecting... Click here if you are not redirected.
2026-09-12 openai: GPT-6 Astra System Card major edit Adds Sept 9 change log; renames 8.7 to Verbalized Metagaming and Oversight Gaming, drops metric plot, adds CoT examples; Alignment section adds generalization caveats.
2026-09-24 openai: GPT-6 Astra System Card dated correction Added approx. 35-page Appendix A covering GPT-6 Sol and Luna. Corrected Astra HealthBench scores after a misconfiguration. Added updated alignment-eval results. Three new changelog entries dated Sept 22.
2026-09-29 google deepmind: Gemini 3.8 Audio (Live, Live Extende... no change log Card expanded to add Gemini 3.8 Flash TTS and Flash-Lite TTS plus Live Avatar video output; added child safety evaluations; frontier safety statement extended; Vertex AI renamed.

What goes quiet

Excluding host-side outage runs and the 86 probes that never reached a server leaves 8,259 valid link checks over 346 documents (33.6 document-years, including documents later removed from the corpus), of which 97.3% succeeded. Hard loss is rare: 2 documents ever returned a 404 (0.06 per document-year). The monitor declared one of them dead after three consecutive 404s, a FAR.AI evaluation of the cyber safeguards of Alibaba’s Qoder agent (Qwen3.8-Max); the other was a slug typo that the publisher fixed with a redirect. Two cases cannot give a reliable rate over 7.3 weeks, but they are not out of line with published baselines for young pages: Pew found 8 % of pages under one year old inaccessible (Pew Research Center 2024), deep links in the Wayback corpus have a median lifetime of 1.3 years (Garg et al. 2024), and 1 to 4 % of web references in scholarly articles from 2012 were already rotten when checked (Klein et al. 2014). Moves are more common: 6 documents moved via permanent redirect (0.18 per document-year: site restructures, a trailing slash, the typo fix and the soft-gone pages below). Bot blocking is the larger nuisance: 16 documents were refused with HTTP 403 at least once, 12 of them OpenAI documents on openai.com. For most of the window the blocks were intermittent (overall a median of 46.1% of later checks), consistent with rate-based bot detection, but in the last week most of the blocked OpenAI pages were refused in most runs; whether that reflects the site or a change on the tracker’s host is not known, and block rates depend on the fetch client’s fingerprint as much as on the site (Gundelach, Mühlhauser, and Herrmann 2026). This matters because OpenAI is also the publisher whose documents carried most of the dated corrections above. The quietest failures were invisible to the link checker: two pages kept returning HTTP 200 while their body became a two-line client-side redirect stub (meta refresh plus a canonical link to the new location), and only the content fingerprint noticed. For at least 17 days the link checker saw nothing, until the publisher replaced the stubs with real permanent redirects on 23 September and the monitor marked them moved (Table 3). This is the reference rot of Zittrain et al. (Zittrain, Albert, and Lessig 2014), the URL resolving while the content is gone, although here it was temporary because the publisher had moved rather than removed the content.

Discovery is fast

For the 80 documents published after the backfill, the median delay from publication to first tracking was 2.3 days and 85.0% were tracked within a week. Leads from index diffs and from the monitor’s own candidate queue were tracked within two days (median 1.3 to 1.5 days); free-form agent search took 4.8 days and manual submissions 14.5 days.

Limitations

The window is 7.3 weeks on one host with one fetch client: enough to measure short-run edit rates, not link rot, whose baselines are annual; block rates in the literature range from under 1 % for browser-like clients to 15 % for headless ones (Gundelach, Mühlhauser, and Herrmann 2026). Coverage was uneven: 41.5% of the daily runs lost their checks to a host-side fault (Appendix B), which stretched the content re-check interval’s tail to 18.0 days at the 90th percentile against an intended 6.7 days, and content changes are observed only when a document’s turn in the 15 % rotation comes up, so change dates lag true edit dates by up to a re-check interval and edits reverted between checks are missed. The instrument changed mid-window (URL migrations and the fingerprint recompute on 31 August). The classification was done by LLM sub-agents with a written taxonomy (189 pairs at high confidence, 48 at medium, 2 at low); it was spot-checked, not double-coded, categories at the boundary of “minor” and “metadata” are soft, and text extraction can drop page elements, so apparent deletions need a check against the raw capture. The corpus is curated by an allow-list and an agent, and the literature baselines cover different populations.

Future work

On the science: a longer horizon allows a hazard model of edit rates by publisher and document type; aligning detected edits with publishers’ change logs and with Watchtower’s disclosed and undisclosed tags would estimate the disclosed fraction; classifying every minted version in the pipeline would turn the appendix into a live feed. On the instrument: re-fetch every document daily instead of a 15 % rotation (the whole corpus is under 400 MB and the link probe already touches every URL), so edits are dated to the day; treat a body that collapses to a redirect stub as a soft 404 and re-point moved documents; keep the new footnote pre-clean under test and diff the raw capture alongside the extracted text so extractor regressions are never read as deletions; extend the furniture patterns to widget rows that end in a score; run on an always-on host.

Every substantive publisher-side change

Table 2 lists all 61 publisher-side version pairs classified as a content change (minor, major, or replaced), in order of detection, with the canonical URL and the stored version ids so the diff can be reproduced from the database (document_versions rows, text under data/text/). Detection dates lag the edit by up to one re-check interval. Table 3 lists the pages that went quiet behind an HTTP 200 and Table 4 the recorded moves.

All substantive publisher-side changes with URL and version ids.
Id Detected Publisher Document and canonical URL Category (flags), versions What changed (classifier summary)
Id Detected Publisher Document and canonical URL Category (flags), versions What changed (classifier summary)
A1 2026-08-17 nvidia Cosmos 3: Omnimodal World Models for Physical AI — Cosmos3-E...
https://huggingface.co/nvidia/Cosmos3-Edge
minor
v55→\tov273
Both code examples change inference defaults: num_inference_steps 50 to 20, guidance_scale 5.0 to 6.0, flow_shift 3.0 to 12.0. Plus download count.
A2 2026-08-19 inclusion ai Ling-3.0-tiny Model Card
https://huggingface.co/inclusionAI/Ling-3.0-tiny
minor (correction)
v219→\tov283
Activated parameter count changed from 1.3B to 1.4B (Non-emb 1.14B) in four places: intro, overview, MoE bullet, evaluation text. Download counter is furniture.
A3 2026-08-19 nvidia NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Model Card
https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
major
v218→\tov282
Added Hardware Matrix, deployment table, SGLang recipes and speculative-decoding flags. Dropped RTX 5090 from supported hardware and removed the runtime thinking-budget claim.
A4 2026-08-20 openai GPT-5.6 — August Updates
https://deploymentsafety.openai.com/gpt-5-6-august-update/gpt-5-6-august-update.pdf
minor (score, safety, correction)
v10→\tov310
Dated Change log added: GPT-5.5 pass@4 on hard-negative protein binding corrected 0.4% to 1.48% (earlier value was pass@1); table updated; rest is reflow.
A5 2026-08-20 nvidia Nemotron-Labs-Audex Model Card (30B-A3B and 2B)
https://huggingface.co/nvidia/Nemotron-Labs-Audex-30B-A3B
minor (correction)
v81→\tov316
Instruct-mode prefix now task-specific (text vs audio); notes the chat template targets audio, so reported text instruct results need a hand-built prompt.
A6 2026-08-20 nvidia NVIDIA NemotronLabs VoiceChat 11B Model Card
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
minor (score, license, correction)
v82→\tov317
The research-purposes-only use statement was deleted and the tool-calling argument accuracy was revised from 44.2% to 42.2%.
A7 2026-08-20 metr Expenditure Horizon: Measuring Optimization Ability, with an...
https://metr.org/blog/2026-07-21-expenditure-horizon/
major (score, safety, scope)
v137→\tov322
Added approx. 200 lines: human returns on NanoGPT (approx. $2,500 per 1%), agent runs for six models with expenditure horizons $600-$3,300, maintainer mergeability review, appendices.
A8 2026-08-20 metr Many SWE-bench-Passing PRs Would Not Be Merged into Main
https://metr.org/notes/2026-03-10-many-swe-bench-passing-prs-would-not-be-merged-into-main/
minor
v145→\tov324
Added Summary paragraph (grader overstates time horizon; trend unsupported) and seven methodological footnotes (Epoch harness lag, 2% corrupted patches, 31 pilot patches, blinding caveats).
A9 2026-08-22 poolside Laguna S 2.1
https://huggingface.co/poolside/Laguna-S-2.1
minor (license)
v254→\tov345
Added paragraph stating Laguna S 2.1 is released under OpenMDW-1.1, fully permissive, with paid support/optimization/indemnification options; counters changed.
A10 2026-08-22 poolside Laguna XS 2.1
https://huggingface.co/poolside/Laguna-XS-2.1
minor (license)
v255→\tov346
Added paragraph stating release under OpenMDW-1.1 (fully permissive, commercial use allowed) with paid support and indemnification options. Counters and widget reorder incidental.
A11 2026-08-22 openai GPT-5.6 System Card
https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf
minor (score, safety, correction)
v7→\tov348
New Aug 19, 2026 change-log entry and table fix: GPT-5.5 pass@4 on hard-negative protein binding corrected 0.4% to 1.5% (was pass@1). Rest is page-number reflow.
A12 2026-08-22 xai Model Card: Grok 4.6
https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf
major (score, safety, correction)
v236→\tov357
Revision 2026-08-17 adds changelog, PartBench, DeepSearchQA, BixBench, KernelBenchInternal v1.1, renumbered sections; corrects HackerBench, self-harm, MASK, LAB scores.
A13 2026-08-24 google deepmind Gemini 3.7 Flash Model Card
https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-7-Flash-Model-Card.pdf
minor (safety)
v277→\tov358
Frontier Safety section changed from ’FSF Report will be published shortly’ to ’is available here’.
A14 2026-08-24 openai GPT-5.6 Preview System Card
https://deploymentsafety.openai.com/gpt-5-6-preview/gpt-5-6-preview.pdf
minor (score, safety, correction)
v214→\tov360
Added dated Change log (Aug 19, 2026) and inline note correcting GPT-5.5 pass@4 on hard-negative protein binding from 0.4% to 1.5%; table updated.
A15 2026-08-25 tencent hunyuan UI-Mate-27B Model Card
https://huggingface.co/tencent/UI-Mate-27B
major (score, scope)
v280→\tov376
Demonstration-guided mode (description, results table, pipeline) removed and assigned to a separate UI-Mate-democua-27B checkpoint in a new checkpoint table.
A16 2026-08-25 tencent hunyuan EVIE-Preview-4.5B Model Card
https://huggingface.co/tencent/EVIE-Preview-4.5B
major (score)
v281→\tov377
Card rewritten around ’Rank #1’ claims: new ViDoRe V3 table with 1,792-token tier (65.36), revised V1+V2 and per-domain scores, Index Cost section.
A17 2026-08-25 nvidia NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Model Card
https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
minor
v282→\tov378
Added one sentence in the training-data section pointing to the Public Summary of Training Content. Download counter and eval-widget rows (Terminal-Bench 2.1 24.58) are furniture.
A18 2026-08-25 inclusion ai Ling-3.0-tiny Model Card
https://huggingface.co/inclusionAI/Ling-3.0-tiny
minor (correction)
v283→\tov379
Activated parameter count reverted from 1.4B (Non-emb 1.14B) back to 1.3B in the same four places. Total 7.9B unchanged.
A19 2026-08-25 openai GPT-5.5 System Card
https://deploymentsafety.openai.com/gpt-5-5/gpt-5-5.pdf
major (score, safety, correction)
v8→\tov387
Change log dated August 19, 2026 added; GPT-5.5 pass@4 on hard-negative protein binding corrected from 0.4% to 1.48%; page numbers shifted.
A20 2026-08-26 nvidia NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning Model Card
https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
minor
v76→\tov398
One line added to the dataset section pointing to a Public Summary of Training Content. Remaining changes are download and Spaces counters.
A21 2026-08-26 nvidia NVIDIA Nemotron Parse 2.0 Model Card
https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0
major (license)
v84→\tov406
License changed from NVIDIA Open Model License Agreement to OpenMDW-1.1; third-party software notice added; Quick Start and vLLM instructions rewritten.
A22 2026-08-27 openai GPT-Rosalind-5.5 System Card
https://deploymentsafety.openai.com/gpt-rosalind-5-5/gpt-rosalind-5-5.pdf
minor (score, safety, correction)
v31→\tov419
Change log (Aug 19, 2026) added: GPT-5.5 pass@4 on hard-negative protein binding corrected from 0.4% to 1.48%; table updated; rest repagination.
A23 2026-08-28 google deepmind Gemini Omni Flash Model Card
https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-Omni-Flash-Model-Card.pdf
minor (scope)
v36→\tov432
Marked Last Updated August 2026; card now also covers Gemini Omni 1.1 Flash; removed paragraph promising T2VA/I2VA/R2VA/editing/image-gen evaluations at API rollout.
A24 2026-09-03 inclusion ai UI-Venus-2-9B
https://huggingface.co/inclusionAI/UI-Venus-2-9B
major (score, license, safety, correction)
v429→\tov479
Result tables rebuilt with baselines and new benchmarks; safety table now OSHarm+OSBlind; Release Status added, license pending; web claim cut to 4,000+ domains.
A25 2026-09-03 google deepmind Gemini 3.6 Flash Model Card
https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-6-Flash-Model-Card.pdf
minor
v14→\tov480
Intended Usage list adds ’complex video reasoning’ and replaces ’multi-week enterprise processes’ with ’enterprise workflows’. No other section changed.
A26 2026-09-03 google deepmind Gemini 3.5 Flash-Lite Model Card
https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-5-Flash-Lite-Model-Card.pdf
minor
v439→\tov481
Benefit and Intended Usage now lists complex video reasoning among use cases; one sentence split. Nothing else changed.
A27 2026-09-03 openai Expanding Daybreak as the Cyber Defense Window Narrows
https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/
minor (safety)
v250→\tov482
Removed the macOS Keychain/Chrome cookies prompt-response comparison table and a customer testimonial. Refusal-rate, ExploitGym, ExploitBench text unchanged.
A28 2026-09-05 inclusion ai Ling-3.0-flash Model Card
https://huggingface.co/inclusionAI/Ling-3.0-flash
minor
v186→\tov491
vLLM install instructions switched from the inclusionAI vllm-ling-v3 fork to upstream vllm-project/vllm. HF use-instructions block, counters and eval-widget rows are furniture.
A29 2026-09-05 nvidia NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 Model Card
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
minor
v270→\tov494
Training-data section gained one line pointing to a Public Summary of Training Content; rest is HF download/Spaces counters and eval-widget reordering.
A30 2026-09-05 nvidia Cosmos 3: Omnimodal World Models for Physical AI — Cosmos3-S...
https://huggingface.co/nvidia/Cosmos3-Super
minor
v54→\tov496
One sentence added to the training-data section pointing to a Public Summary of Training Content; download counter changed.
A31 2026-09-05 nvidia NVIDIA-Nemotron-3-Super-120B-A12B-BF16 Model Card
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
minor
v53→\tov495
Spec table gains Speculative Decoding row (MTP head, MTPv2 separate checkpoint); training section links NVIDIA’s Public Summary of Training Content.
A32 2026-09-05 nvidia Cosmos 3: Omnimodal World Models for Physical AI — Cosmos3-E...
https://huggingface.co/nvidia/Cosmos3-Edge
minor
v273→\tov497
Added ’Update August 25, 2026’ notice that checkpoint, runtime defaults, examples and benchmarks were updated; link to Public Summary of Training Content.
A33 2026-09-05 openai OpenAI and Hugging Face partner to address security incident...
https://openai.com/index/hugging-face-model-evaluation-security-incident/
minor (safety)
v104→\tov499
Pending-review sentence replaced by an ’Update on August 26, 2026’ linking to published findings on the Hugging Face incident.
A34 2026-09-06 palisade research Language Models Can Autonomously Hack and Self-Replicate
https://palisaderesearch.org/blog/self-replication
replaced
v194→\tov508
Blog URL now serves a meta-refresh stub ’Redirecting... Click here’ pointing to palisaderesearch.org/research/self-replication; article text no longer captured.
A35 2026-09-06 palisade research Technical Report: Shutdown Resistance in Large Language Mode...
https://palisaderesearch.org/blog/shutdown-resistance-on-robots
replaced (safety)
v195→\tov509
Page body (shutdown-resistance-on-a-robot report: 3/10 physical, 52/100 simulated trials) replaced by a two-line redirect stub: Redirecting... Click here if you are not redirected.
A36 2026-09-06 google deepmind Gemini 3.7 Flash Model Card
https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-7-Flash-Model-Card.pdf
minor
v441→\tov513
Description adds ’support for agentic video understanding’; intended use cases add ’complex video reasoning’.
A37 2026-09-09 nvidia GR00T-H model card
https://huggingface.co/nvidia/GR00T-H
minor (license)
v328→\tov525
Header notice now labels this the N1.6 non-commercial version and points to GR00T-H-N1.7 as the commercially usable version.
A38 2026-09-10 tencent hunyuan Hy4 preview Model Card
https://huggingface.co/tencent/Hy4-preview
minor
v433→\tov533
Deployment section dropped build-vLLM-from-source steps and the explicit vllm serve MTP command, pointing to recipes and the prebuilt image instead.
A39 2026-09-12 openai GPT-6 Astra System Card
https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf
major (safety)
v500→\tov551
Adds Sept 9 change log; renames 8.7 to Verbalized Metagaming and Oversight Gaming, drops metric plot, adds CoT examples; Alignment section adds generalization caveats.
A40 2026-09-14 openai Strengthening societal resilience with Rosalind Biodefense
https://openai.com/index/strengthening-societal-resilience-with-rosalind-biodefense/
minor (safety)
v462→\tov561
The paragraph naming one launch partner in the biodefense access program, and that partner’s quote, were deleted. Nothing added.
A41 2026-09-14 openai Introducing GPT-Rosalind for life sciences research
https://openai.com/index/introducing-gpt-rosalind/
major (license, safety)
v465→\tov562
Sept 11, 2026 update: exits research preview, global trusted access, pricing from Oct 5. Removed free-credits paragraph, customer quote, and preview-terms compliance clause.
A42 2026-09-14 openai OpenAI Daybreak: Trusted Access for Cyber — Overview
https://help.openai.com/en/articles/20001258-openai-daybreak-trusted-access-for-cyber-overview
minor (safety)
v466→\tov563
Two FAQ entries added: reduced refusals on Astra only for Daybreak Red, not Blue; Daybreak Red is org-only and not open to individual applicants.
A43 2026-09-15 deepseek DeepSeek-V4-Flash-Vision-Exp
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
minor
v473→\tov564
Added a ’How to Run with vLLM’ section with a 4xGB300 docker serving command using DSpark speculative decoding. Widget and nav changes are furniture.
A44 2026-09-16 anthropic An alignment assessment of recent cybersecurity incidents
https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
minor (safety, correction)
v540→\tov579
Sept 10 correction note added; PyPI removal window changed from approx. 90 minutes to under an hour; research model penetrated one neighbor, not several. Related-content footer changed.
A45 2026-09-16 deepseek DeepSeek-V4.1-Flash
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
minor
v541→\tov580
Agentic evaluation methodology note rewritten: per-benchmark harness assignments, no-network condition for one benchmark, sampling settings stated. No score changed.
A46 2026-09-16 anthropic Real-time cyber safeguards on Claude Opus and Sonnet
https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet
major (safety, scope)
v542→\tov581
Cyber Verification Program now available via ’Claude on Google Cloud’ (was unavailable on Vertex); new section requires consenting to the Advanced AI Safety Addendum and setting PublisherModelConfig data-retention fields per model.
A47 2026-09-16 far ai Revisiting Frontier LLMs’ Attempts to Persuade on Extreme To...
https://www.far.ai/blog/revisiting-attempts-to-persuade
major (score, safety, scope)
v193→\tov583
Two sections added: reasoning-level effect (Gemini 3 harmful persuasion approx. 90% low vs approx. 25% high; GPT-5.1 and Opus 4.5 approx. 0%) and Gemini 3 Flash results.
A48 2026-09-23 anthropic Introducing Claude Fable 5.1 and Claude Mythos 5.1
https://www.anthropic.com/claude-fable-and-mythos-5-1
minor (safety)
v488→\tov605
Added methodology note: Fable scored zero on OSWorld/AutomationBench where production safeguards intervened; other cyber/bio interventions completed by Opus 4.8/Opus 5. Minor table cell formatting.
A49 2026-09-23 stepfun StepAudio 3 Realtime Technical Report
https://arxiv.org/abs/2609.14005
major
v586→\tov608
arXiv revision v2 (19 Sep 2026) posted with a submission history added. The abstract and its scores are unchanged; the PDF shrank from 937 KB to 367 KB.
A50 2026-09-23 nvidia Cosmos 3: Omnimodal World Models for Physical AI — Cosmos3-S...
https://huggingface.co/nvidia/Cosmos3-Super
major (scope)
v496→\tov613
Removed Cosmos3-Nano, Nano-Policy-DROID, Super-Image2Video and Super-Text2Image from Model Versions and parameter lists; card now covers Cosmos3-Super (64B) only.
A51 2026-09-23 nvidia Cosmos 3: Omnimodal World Models for Physical AI — Cosmos3-E...
https://huggingface.co/nvidia/Cosmos3-Edge
major (scope)
v497→\tov614
Removed eight other Cosmos3 variants (Edge-Policy-DROID, Super 4-step distills, 05/31 Nano/Super release block) from model versions and parameter lists; card now Cosmos3-Edge only.
A52 2026-09-24 openai GPT-6 Astra System Card
https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf
major (score, safety, scope, correction)
v551→\tov620
Added approx. 35-page Appendix A covering GPT-6 Sol and Luna. Corrected Astra HealthBench scores after a misconfiguration. Added updated alignment-eval results. Three new changelog entries dated Sept 22.
A53 2026-09-25 inclusion ai Ling-3.0-tiny Model Card
https://huggingface.co/inclusionAI/Ling-3.0-tiny
minor
v379→\tov634
vLLM install now uses upstream vllm-project/vllm instead of inclusionAI fork; added ’Training content summary’ section pointing to training-content disclosure document.
A54 2026-09-26 openai OpenAI Daybreak: Trusted Access for Cyber — Overview
https://help.openai.com/en/articles/20001258-openai-daybreak-trusted-access-for-cyber-overview
major (safety, scope)
v563→\tov646
Tier table rebuilt: Blue gives reduced refusals on GPT-5.5/5.6 Sol/GPT-6 Sol/Luna; Red adds GPT-5.5-Cyber and Astra; GPT-5.6-Cyber needs extra approval; FIDO2 key requirement added.
A55 2026-09-28 inclusion ai Ling-2.6-1T model card
https://huggingface.co/inclusionAI/Ling-2.6-1T
minor
v530→\tov659
Added short Training content summary section linking public training-content disclosure; HF download counter changed.
A56 2026-09-28 inclusion ai Ling-2.6-flash model card
https://huggingface.co/inclusionAI/Ling-2.6-flash
minor
v531→\tov660
Added ’Training content summary’ section pointing to a public training-content disclosure document; states it does not replace docs, usage terms or license.
A57 2026-09-28 inclusion ai Ring-2.6-1T model card
https://huggingface.co/inclusionAI/Ring-2.6-1T
minor
v532→\tov661
Added short Training content summary section linking public training-content disclosure; other changes are HF download count and eval-widget row reorder.
A58 2026-09-28 inclusion ai Ling-3.0-flash-VL
https://huggingface.co/inclusionAI/Ling-3.0-flash-VL
minor (correction)
v528→\tov662
Context window corrected from 1M to 256K tokens in two places; dropped –mem-fraction-static serving flag; added Training content summary section pointing to disclosure document.
A59 2026-09-29 anthropic Real-time cyber safeguards on Claude Opus and Sonnet
https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet
minor (safety, scope)
v581→\tov671
Scope note now excludes Claude Opus 5.5 and Sonnet 5.5 from these cyber safeguards; says Cyber Verification Program will expand to Opus/Sonnet 5.5 and Mythos.
A60 2026-09-29 google deepmind Gemini 3.8 Audio (Live, Live Extended Thinking) Model Card
https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-8-Audio-Model-Card.pdf
major (safety, scope)
v584→\tov677
Card expanded to add Gemini 3.8 Flash TTS and Flash-Lite TTS plus Live Avatar video output; added child safety evaluations; frontier safety statement extended; Vertex AI renamed.
A61 2026-09-29 redwood research Astra is much better at reasoning with filler tokens than pr...
https://blog.redwoodresearch.org/p/astra-is-much-better-at-reasoning
minor
v623→\tov686
Added a 10-shot prompting filler-token result: Astra improves significantly and other models do not. The table is linked to LessWrong, not included.
Pages whose body became a redirect stub while the URL kept returning HTTP 200, and when the tracker first recorded a real redirect (moved) or a 404 (dead) for them.
Id Detected Document URL that kept returning HTTP 200 Resolved
Id Detected Document URL that kept returning HTTP 200 Resolved
S1 2026-09-06 palisade-research-gpt-5-4-independent-eval https://palisaderesearch.org/blog/self-replication moved 2026-09-23
S2 2026-09-06 palisade-research-grok-4-grok-4-0709-independent-eval https://palisaderesearch.org/blog/shutdown-resistance-on-robots moved 2026-09-23

Tracker reliability

Of the 53 runs, 22 recorded a fetch error for every URL, the longest streak lasting 4 consecutive runs (Figure 2); the last was on 22 September, and none of the 9 runs since then was a full outage. The snapshot does not record why the outages stopped, and the run logs stop on the same date. The systemd journal explains 21 of them: the run started within two minutes of the laptop waking from suspend (median 0 s after resume, versus 38,412 s for healthy runs). The timer is Persistent=true, so a trigger missed while asleep fires at the moment of resume, before Wi-Fi is back, and every probe fails instantly. The run then continues; the agent phase, minutes later, usually has network again (9 outage runs still added documents) and recent runs label the commit as an outage, but that day’s link and content checks are lost. A further two runs lost part of their checks (86 of 287 probes on 18 September, while the network flapped). The curation agent failed in 16 runs, 13 of them outage runs where its credential refresh had no network either; only 3 healthy runs lost their agent, and 19 of the 44 runs with a run log had both phases healthy. Run logs after 22 September were not in the snapshot, so the agent phase of the last 9 runs is unknown (“no log” in Table 5). Outages stretched the tail of the content re-check interval. Table 5 gives every run.

Link-check outcome of every tracked document (rows, grouped by publisher, ordered by date first tracked; the largest publishers labelled where space allows) in every run (columns). Grey columns: every check failed on the tracker’s side; light blue: reachable; orange: HTTP 403; red: 404; violet: permanent redirect; amber: fetch error in a healthy run; white: not yet tracked.
Every daily run: link probes and their outcomes, content checks, whether the run was a full outage, seconds between the host’s last resume from suspend and the run start, and the agent phase.
Run Links Reachable Blocked Errors Content checks Changed Outage Resume (s) Agent
Run Links Reachable Blocked Errors Content checks Changed Outage Resume (s) Agent
2026-08-11 06:59 193 0 0 193 29 0 yes 0 failed
2026-08-12 06:17 193 190 3 0 29 14 67309 ok
2026-08-13 06:20 199 196 3 0 30 15 1217 ok
2026-08-14 07:00 203 0 0 203 31 0 yes 0 ok
2026-08-15 08:22 215 0 0 215 33 0 yes 0 failed
2026-08-16 13:21 222 0 0 222 34 0 yes 0 failed
2026-08-17 06:15 222 222 0 0 34 18 6674 ok
2026-08-18 07:25 225 0 0 225 34 0 yes 0 ok
2026-08-19 06:17 227 227 0 0 35 21 47075 ok
2026-08-20 06:18 233 229 4 0 35 18 1421 ok
2026-08-21 07:27 237 0 0 237 36 0 yes 0 ok
2026-08-22 06:17 245 245 0 0 37 16 48912 ok
2026-08-23 06:22 247 0 0 247 38 0 yes 0 ok
2026-08-24 06:20 247 245 2 0 38 14 659 failed
2026-08-25 06:17 247 247 0 0 38 26 44312 ok
2026-08-26 06:50 249 249 0 0 38 19 49738 ok
2026-08-27 06:17 254 254 0 0 39 14 799 ok
2026-08-28 07:28 260 0 0 260 43 0 yes 0 ok
2026-08-29 09:09 261 0 0 261 40 0 yes 0 failed
2026-08-30 09:37 261 0 0 261 40 0 yes 0 ok
2026-08-31 07:11 262 0 0 262 40 0 yes 0 failed
2026-09-01 06:18 241 241 0 0 37 6 46045 ok
2026-09-02 07:03 246 0 0 246 37 0 yes 0 failed
2026-09-03 06:16 246 245 1 0 37 10 1319 failed
2026-09-04 07:26 246 0 0 246 37 0 yes 0 failed
2026-09-05 06:19 250 250 0 0 38 9 47587 ok
2026-09-06 06:15 258 258 0 0 39 7 133747 ok
2026-09-07 06:24 264 0 0 264 40 0 yes 0 ok
2026-09-08 07:14 266 0 0 266 40 0 yes 0 failed
2026-09-09 06:17 266 254 12 0 40 3 47982 ok
2026-09-10 06:19 269 269 0 0 41 11 69309 ok
2026-09-11 06:16 272 272 0 0 41 5 2744 ok
2026-09-12 06:17 275 271 3 0 42 2 1534 ok
2026-09-13 07:04 277 0 0 277 42 0 yes 0 ok
2026-09-14 06:17 279 278 0 0 42 7 591 ok
2026-09-15 08:12 279 266 12 1 42 5 0 ok
2026-09-16 10:20 282 0 0 282 43 0 yes 0 failed
2026-09-16 12:52 282 282 0 0 43 12 9134 ok
2026-09-17 09:08 287 0 0 287 44 0 yes 0 failed
2026-09-18 06:17 287 187 14 86 44 2 1240 failed
2026-09-19 09:15 287 0 0 287 44 0 yes 0 failed
2026-09-20 06:18 287 0 0 287 44 0 yes 49325 failed
2026-09-21 06:51 287 0 0 287 44 0 yes 0 ok
2026-09-22 08:07 290 0 0 290 44 0 yes 0 failed
2026-09-23 06:15 302 299 3 0 46 10 45633 no log
2026-09-24 06:17 305 295 9 0 46 5 76975 no log
2026-09-25 06:17 310 298 11 0 47 10 38412 no log
2026-09-26 06:17 318 304 13 0 48 6 124812 no log
2026-09-27 11:24 318 304 13 0 48 5 0 no log
2026-09-28 06:16 318 306 11 0 48 10 35920 no log
2026-09-29 06:19 318 317 0 0 48 13 40709 no log
2026-09-29 06:52 318 305 12 0 48 5 42713 no log
2026-09-29 07:44 318 317 0 0 48 15 1540 no log
Garg, Kritika, Sawood Alam, Daniel Ayala, Michael L. Nelson, and Michele C. Weigle. 2024. “Some URLs Are Immortal, Most Are Ephemeral.” Web Science and Digital Libraries Research Group blog, Old Dominion University.
Gundelach, Robin, Michael Mühlhauser, and Dominik Herrmann. 2026. “Detecting Bot Detection: Prevalence, Techniques, and Implications for Web Measurement Research.” arXiv:2606.14525.
Klein, Martin, Herbert Van de Sompel, Robert Sanderson, Harihar Shankar, Lyudmila Balakireva, Ke Zhou, and Richard Tobin. 2014. “Scholarly Context Not Found: One in Five Articles Suffers from Reference Rot.” PLOS ONE 9 (12): e115253. https://doi.org/10.1371/journal.pone.0115253.
Pew Research Center. 2024. “When Online Content Disappears.” Pew Research Center.
The Midas Project. 2026. “AI Safety Watchtower.” themidasproject.com/watchtower, accessed 2026-09-22.
Zittrain, Jonathan, Kendra Albert, and Lawrence Lessig. 2014. “Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations.” Harvard Law Review Forum 127: 176.

  1. Every number, table and figure is generated by analyze.py from a read-only snapshot of the cardtrack database. Version pairs were classified by nine Claude Fable 5.1 sub-agents and spot-checked by hand. Appendix A lists every substantive change with its URL and version ids.↩︎