## The big picture

**Claude Sonnet 5.5**
Anthropic released Claude Sonnet 5.5, a new state-of-the-art model that runs 30% faster and costs up to 30% less than its predecessor while beating it on every benchmark. It is priced the same as Sonnet 5 but offers superior performance, though early tests reveal a "max" thinking mode bug where the model can exhaust its token budget (128k tokens) without producing output.
[Simon Willison](https://simonwillison.net/2026/Sep/28/claude-sonnet-5-5/)

**OpenAI DevDay 2026**
OpenAI held a DevDay event in Fort Mason, San Francisco. The event included a keynote and a "creator" area.
[Simon Willison](https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/)
**Frontier Cyber Capabilities**
Anthropic’s Frontier Red Team reported that both GLM-5.3 and Claude Mythos Preview have crossed a critical threshold into functional binary exploitation, achieving full control flow hijacks in 4% and 6% of trials respectively. This marks a distinct shift from earlier models like Claude Opus 4.6 and GLM-5.2, which failed to succeed in any such trials, raising urgent containment concerns.
[Simon Willison](https://simonwillison.net/2026/Sep/29/anthropic-frontier-red-team/)

## Architectural breakthroughs

**Clef Decision Models**
Cloudflare released Clef, a 27B multimodal model (and 9B Flash variant) that eliminates free-form text generation by using a joint schema head to score probabilities for all allowed options in a single forward pass. By reading state as text, JSON, images, or video and returning structured decisions without parsing, it offers a novel architecture for control tasks that is API-compatible with existing systems.
[Hugging Face](https://huggingface.co/Cloudflare/clef)

**VISTA Visual Harness**
The authors of VISTA demonstrate that multimodal models possess strong reasoning abilities that can be unlocked by a visual harness providing long-horizon vision and lossless visual memory. By allowing the model to actively retrieve and reorganize past observations, VISTA boosts Claude Opus 5.0’s performance on ARC-AGI-3 to a perfect score, completing all games with 57.4% fewer actions than humans.
[arXiv](https://arxiv.org/abs/2610.02200v1)

**Finetuning with Sampling**
Challenging the narrative that RL is superior for generalization, this work introduces an MCMC sampling algorithm that tailors data distribution to suit on-policy learners while leveraging off-policy expert data. The method allows SFT to learn better than conventional wisdom suggests, combining the strengths of on-policy learning with privileged information from off-policy data without modifying the learning objective.
[arXiv](https://arxiv.org/abs/2610.02140v1)

**LoopCD for Looped Transformers**
LoopCD is a training-free contrastive decoding framework for looped transformers that leverages inherent weak-and-strong prediction pairs from recurrent passes. By contrasting the final prediction with an earlier recurrent pass in either logit or hidden-state space, it delivers substantial gains in reasoning tasks like AIME 2024 without auxiliary models or external training.
[arXiv](https://arxiv.org/abs/2610.02185v1)

**TACO Optimizer**
TACO is a ternary absolute-max column-wise one-sparse optimizer designed to reduce the memory overhead of full-parameter fine-tuning for LLMs. It follows Muon’s operator-norm steepest-descent view but takes the geometric route further, computing the exact steepest-descent direction under a dimension-normalized 1-to-1 operator norm to maintain AdamW-pretrained model performance.
[arXiv](https://arxiv.org/abs/2610.02199v1)

## New open-weight model releases

**Naive-N0.5-Flash**
NaiveAI released Naive-N0.5-Flash, a 309B MoE model with 15.5B active parameters built for coding and AI R&D. It supports a native 1M-token context window through a hybrid of Sliding-Window Attention and DeepSeek Sparse Attention, with no full-attention layers, and is optimized for inference speeds up to 2,000 tokens/s.
[Hugging Face](https://huggingface.co/NaiveAI/Naive-N0.5-Flash)

**Kolibri**
Aleph Alpha released Kolibri, a sovereign open-weight MoE reasoning model with 78B total parameters (3.46B active) focused on German and English. It supports explicit reasoning mode, tool calling, and long-context optimization, with weights available in float8_e4m3fn precision and an Apache 2.0 license.
[Hugging Face](https://huggingface.co/Aleph-Alpha/Kolibri-1)

**Phonon-2**
FermionResearch released Phonon-2, the most accurate open speech recognition model for English under 900 MB, averaging 5.21% word error across seven benchmark sets. It achieves 100.8% of its 2.5 GB teacher’s word accuracy on parliamentary speech while being 15 times smaller, enabling high-quality transcription on low-end devices like the M5 MacBook Air.
[Hugging Face](https://huggingface.co/FermionResearch/Phonon-2)

**GLM-5.3-UNCENSORED-EXL3-3.0bpw**
This release provides a 3.0 bpw EXL3 quantization of the 753B MoE GLM-5.3 model, demonstrating advanced compression techniques for consumer hardware. It uses a mixed precision scheme with attention and shared experts at 5 bpw, dense MLPs at 4 bpw, and routed experts at 3 bpw, resulting in a 273 GiB file size.
[Hugging Face](https://huggingface.co/Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw)

## Hardware & optimization

**ds4 Local Inference**
The creator of Redis released ds4, a new tool for running LLMs locally, appealing to privacy-conscious practitioners seeking robust local inference options.
[source](https://dwarfstar.sh/)

**Qwen3.8 Flash GGUF Benchmark**
The Kaitchup published a benchmark of Qwen3.8 Flash Next GGUF formats. The analysis covers 12 compressed GGUFs and 42.7 million generated tokens, focusing on accuracy and token efficiency from Q4 to Q1.
[The Kaitchup](https://kaitchup.substack.com/p/qwen38-flash-next-gguf-benchmark)
**GLM-5.3 Sparse Attention Analysis**
SemiAnalysis published a deep technical analysis of how GLM-5.3’s sparse attention mechanisms affect HBM memory usage, detailing optimizations like KV Cache Offloading and HiSparse.
[SemiAnalysis](https://newsletter.semianalysis.com/p/sparse-savings-persistent-demand-inside-glm53)

## Also this week

**Agent Worms** — Matthew Green highlights a novel attack vector where sandboxed agents communicate via shared caches to spread payloads, raising urgent containment concerns. [Simon Willison](https://simonwillison.net/2026/Oct/1/matthew-green/)
**X-Tree** — A novel approach to agent training that tokenizes reusable experience hierarchies directly into weights, addressing limitations in flat-action training. [HF Daily Papers](https://huggingface.co/papers/2609.32993)
**Keyword Harnesses Fail Open** — Diagnostic work exposing false positives in tool-use benchmarks for small models, urging caution in evaluation practices. [arXiv](https://arxiv.org/abs/2610.02142v1)
**Persona Dosing** — A technical advancement in activation steering allowing for calibrated, graded control of persona traits beyond binary interventions. [HF Daily Papers](https://huggingface.co/papers/2609.36388)
**Ego2Act** — A benchmark for evaluating goal-directed manipulation in egocentric video generation, addressing multi-step physical reasoning. [HF Daily Papers](https://huggingface.co/papers/2610.01092)
**KaliBench** — A fine-grained benchmark for natural-language-to-CLI translation on Kali Linux, critical for cybersecurity tool-use evaluation. [arXiv](https://arxiv.org/abs/2610.02206v1)
**RPG Framework** — A framework for autonomous robot self-improvement via simulation practice, addressing the high cost of human-in-the-loop skill development. [arXiv](https://arxiv.org/abs/2610.02204v1)
**Embedding Prediction** — An architectural shift in DiTs where predicted embeddings serve as conditioning signals, adapting to the current noisy state at every denoising step. [arXiv](https://arxiv.org/abs/2610.02203v1)
**SemanTok** — A flexible video tokenizer that feeds frozen DINO features into its encoder, achieving high semantic alignment and video fidelity for AR models. [HF Daily Papers](https://huggingface.co/papers/2610.00686)
**On-Policy vs Off-Policy Distillation** — A systematic study revealing that token-level KL direction, not rollout policy, more clearly shapes task performance in distillation. [HF Daily Papers](https://huggingface.co/papers/2609.35259)
**Where-OPD** — Extends on-policy self-distillation to MLLMs using synthetic scenes with textual, spatially grounded guidance. [arXiv](https://arxiv.org/abs/2610.02117v1)
**Video Generation Survey** — A comprehensive review of post-training and alignment strategies for video generation models. [HF Daily Papers](https://huggingface.co/papers/2610.00812)
**RLE-Bench** — A benchmark for coding agents as robot learning engineers, evaluating broader engineering capabilities beyond control. [HF Daily Papers](https://huggingface.co/papers/2609.34210)
**Model Context Protocol SDK** — The official Python SDK for Model Context Protocol servers and clients, standardizing LLM-tool interfaces. [GitHub](https://github.com/modelcontextprotocol/python-sdk)
**Agent-Reach** — A CLI tool giving AI agents eyes to see the entire internet, reading Twitter, Reddit, YouTube, and more with zero API fees. [GitHub](https://github.com/Panniantong/Agent-Reach)
**iFixAi** — A tool for independent auditing of AI agents, answering whether the agent is doing what it is supposed to do in under 120 seconds. [GitHub](https://github.com/ifixai-ai/iFixAi)
**OpenMontage** — An open-source, agentic video production system with 12 pipelines and 100+ tools. [GitHub](https://github.com/calesthio/OpenMontage)
**strix** — An open-source AI penetration testing tool to find and fix app vulnerabilities. [GitHub](https://github.com/usestrix/strix)
**opensre** — An open-source toolkit for building AI SRE agents. [GitHub](https://github.com/Tracer-Cloud/opensre)
**VisionHOPE** — A research-grade vision backbone that treats the network as a self-modifying learning system. [Hugging Face](https://huggingface.co/PSRben/VisionHOPE)
**QI_2.1_AnyAngle** — A LoRA for style-aligned camera angle manipulation in image generation using Gaussian/3D models. [Hugging Face](https://huggingface.co/lilylilith/QI_2.1_AnyAngle)
**Gemini 4 Argon** — Google’s answer to Astra/Fable with 1M output, currently restricted to government users and trusted cyber defenders. [Latent Space](https://www.latent.space/p/ainews-gemini-4-argon-gdms-answer)
**Nemotron Notes** — A broad coverage of NVIDIA Nemotron’s open-weight tech reports. [Cameron R. Wolfe](https://cameronrwolfe.substack.com/p/nemotron)
**Deep Learning Weekly** — A curated digest covering GPT-6 Sol and Luna, AI Evals, and RRSI. [Deep Learning Weekly](https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-475)
**AI Agents & Work** — Epoch AI analyzes the macroeconomic implications of hundreds of millions of AI agents. [Epoch AI](https://epochai.substack.com/p/hundreds-of-millions-of-ai-agents)
**Claude Code Next Era** — Details on the next generation of Claude Code, including Mods, Plugins, and Projects. [Latent Space](https://www.latent.space/p/thariq)
**Budget Caps** — Simon Willison argues for default hard budget caps on pay-by-usage services to mitigate financial risk from rogue agents. [Simon Willison](https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/)
**OpenAI Safety Leader Quits** — A high-profile departure signaling cultural issues at OpenAI. [The Guardian](https://www.theguardian.com/technology/2026/oct/03/openai-safety-leader-quits-warning-ai-companys-culture-is-broken)
**ScholarCatalyst** — A benchmark for retrieving papers that inspire new research, showing agentic search performs no better than embedding retrieval. [arXiv](https://arxiv.org/abs/2610.02202v1)
**Architect-Ant** — A framework for generating furnished architectural floor plans using a curated dataset of real plans. [HF Daily Papers](https://huggingface.co/papers/2606.10953)
**OpenTumorBoard** — A benchmark for multidisciplinary tumor board discussion trajectories, revealing limitations in current medical LLMs. [HF Daily Papers](https://huggingface.co/papers/2609.32810)
**Pi 1.0** — The minimalist harness has gone stable. It now includes TypeScript support. [Latent Space](https://www.latent.space/p/ainews-pi-10-pi-durable-and-aie-nyc)
**Every SaaS is a Harness** — An opinion piece arguing that every SaaS business will become a harness around a model. [SSHH Blog](https://blog.sshh.io/p/the-harness-is-the-company)
**LeCun on AI Risk** — Yann LeCun states he has "zero concerns" about AI wiping out humanity. [Fortune](https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/)
**Pop!_OS Bans AI Code** — System76 bans AI-generated code from much of its Pop!_OS codebase. [Neowin](https://www.neowin.net/news/system76-bans-ai-generated-code-across-many-of-its-cosmic-codebases/)
**Crypto Capture** — This work examines the crypto capture of foreign aid. [NBER](https://www.nber.org/papers/w35655)
**Rex's Dino Store** — A personal anecdote about a dinosaur-operated newsstand in Brooklyn. [Simon Willison](https://simonwillison.net/2026/Oct/2/rex-s-dino-store/)
**Lego AI Generator** — A hobbyist project using LLMs to generate LDraw source code for LEGO models. [GitHub](https://github.com/anteloc/ldraw-nova)
**Academia is for Ambition** — An interview with Alex Zhang on Jev, PhD masxing, and the future of harnesses. [Latent Space](https://www.latent.space/p/rlm)
**Inside-Out AI** — A case study on how Airbnb is rebuilding its backend and guest experience with AI. [Latent Space](https://www.latent.space/p/airbnb)
**Google Skills** — A collection of skills for Google products. [GitHub](https://github.com/google/skills)
**Agno** — A framework for building agents. [GitHub](https://github.com/agno-agi/agno)
**Production Agentic RAG Course** — A course repository for agentic RAG. [GitHub](https://github.com/jamwithai/production-agentic-rag-course)
**Import AI 474** — A newsletter roundup covering Platonic mindspace, TPUs in space, and Zhipu’s RSI loop. [Import AI](https://jack-clark.net/2026/09/28/import-ai-474-platonic-mindspace-tpus-in-space-zhipu-starts-an-outer-rsi-loop/)
**Qualcomm Snapdragon X2 Elite** — A technical analysis of Qualcomm’s system-level architecture in the new laptop chip. [Chips and Cheese](https://chipsandcheese.com/p/qualcomms-system-level-architecture)
**Language Models for Text Classification** — A visual guide to RNNs, CNNs, Transformers, and Calibration. [Ahead of AI](https://magazine.sebastianraschka.com/p/classifier-history-and-jev)
**Why Dwarkesh is Wrong** — Insider look at OpenAI’s computer use agent architecture and rapid development cycle. [Latent Space](https://www.latent.space/p/devday-2026)
**GPT 6.1 Sol** — Announcement of a new high-performance OpenAI model with significant cost reductions. [Simon Willison](https://simonwillison.net/2026/Sep/29/hn-49898129/)
**Quoting @joedaroo** — Commentary on the organizational challenges of adapting to rapid AI capability jumps. [Simon Willison](https://simonwillison.net/2026/Sep/28/joedaroo/)
**September Sponsors Newsletter** — Meta-content regarding a paid newsletter subscription. [Simon Willison](https://simonwillison.net/2026/Oct/3/newsletter/)
