## The big picture

**Anthropic and OpenAI release new models**
Anthropic released Claude Opus 5.5, while OpenAI released more efficient GPT-6 models. Prices have been cut by 40-50%.
[Latent Space](https://www.latent.space/p/ainews-claude-opus-55-the-new-default)
**Meta’s Muse brings persistent Linux VMs to consumers**
Meta’s new agentic AI system, Muse, provides each user with a persistent Linux virtual machine in the cloud, packaged with a consumer-friendly mascot interface. While technically groundbreaking for its stateful, long-horizon execution, experts warn that users may underestimate the power and potential dangers of such unrestricted agentic access.
[Simon Willison](https://simonwillison.net/2026/Sep/25/john-gruber/)

## Architectural breakthroughs

**Rufus-Air: A reproducible 106B post-training recipe**
The authors release a fully documented, eight-stage post-training pipeline for a 106B model, progressing from SFT to specialized RL for coding, agents, and search. The work demonstrates that diverse, high-quality SFT establishes a strong capability floor, while difficulty filtering and reward reliability are critical for ordering RL stages without new human annotation.
[HF Daily Papers](https://huggingface.co/papers/2609.29421)

**Superposition Linearity in Transformers**
Researchers provide evidence that Transformers exhibit fundamental linearity, where linearly combined inputs yield superposed next-token distributions, a property intrinsic to the architecture rather than emergent from training. They show this linearity diminishes during pretraining but can be restored via lightweight fine-tuning, enabling a new guided decoding procedure to disentangle superposed outputs.
[HF Daily Papers](https://huggingface.co/papers/2609.29845)

**Agent-Editing World Models**
To address state contamination in long-horizon agents, the authors propose an Agent-Editing World Model that predicts task progress rather than simulating tool responses. It combines an Action Judge to classify decisions with State Revision to edit noisy reasoning continuations, significantly improving agent performance by preventing outdated plans from distorting subsequent actions.
[HF Daily Papers](https://huggingface.co/papers/2609.28416)

**Neural Spectral Capacity**
The authors introduce Neural Spectral Capacity, a closed-form scalar derived from the singular-value spectrum of weight matrices, allowing architecture evaluation from specification alone without instantiation. This metric enables an exact dynamic-programming solver to find globally optimal architectures under resource constraints, outperforming black-box search methods.
[HF Daily Papers](https://huggingface.co/papers/2609.23087)

## New open-weight model releases

**MiMo-V2.6-Pro-RL**
Xiaomi releases a flagship omnimodal model trained with scaled reinforcement learning for self-improvement, supporting text, image, video, and audio with 1M token context. The model uses a single mixed RL run across coding, agents, and cybersecurity, demonstrating expanded capability frontiers through exploration and feedback.
[Hugging Face](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL)

**Qwen-Image-2.1-viggle-turbo**
Viggle releases a distilled student of Qwen-Image-2.1 that achieves 5x speedup by reducing transformer passes from 40 to 6 with no classifier-free guidance. The model maintains competitive quality on most prompts, with minor gaps in dense text rendering, enabling faster local and cloud image generation.
[Hugging Face](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo)

**LensVLM-9B**
Apple releases a 9B Vision Language Model that scans compressed images of text and selectively expands only relevant pages to uncompressed form via learned tools. This selective context expansion mechanism improves efficiency for document-heavy tasks while maintaining high accuracy.
[Hugging Face](https://huggingface.co/apple/LensVLM-9B)

**CLM-v0.1-8B**
This 8B contrastive language model uses a bidirectional InfoNCE loss to connect states and actions, achieving latency up to 9x lower than Jev on computer-use and tool-calling tasks. Fine-tuned as a verifier, it sets new state-of-the-art results on DeepSWE and Terminal-Bench while being significantly faster.
[Hugging Face](https://huggingface.co/Contrastive-LM/CLM-v0.1-8B)

## Hardware & optimization

**NVIDIA Model-Optimizer**
NVIDIA releases a unified library for SOTA model optimization techniques, including quantization, distillation, pruning, and speculative decoding. The tool compresses models for downstream deployment in TensorRT-LLM and vLLM, significantly lowering the barrier for production inference optimization.
[GitHub](https://github.com/NVIDIA/Model-Optimizer)

**Soup: Layer streaming for 4GB VRAM**
The authors demonstrate training an 8B model on a 4GB laptop GPU using layer streaming, a technique that loads model layers sequentially to minimize memory footprint. This approach makes fine-tuning large models accessible on consumer hardware with a single YAML configuration.
[GitHub](https://github.com/MakazhanAlpamys/Soup)

**Qwen-Image-2.1 GGUF quantizations**
GGUF quantizations of Qwen-Image-2.1 are available for local image generation, including uncensored variants. These releases utilize the original upstream base weights.
[Hugging Face](https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF)
## Also this week

- **Revealing the details of how OpenAI agents hacked Hugging Face** — [source](https://swarmtraces.org/)
- **Hindsight** — A novel learning-based memory mechanism for long-horizon agents. [GitHub](https://github.com/vectorize-io/hindsight)
- **Training Object Permanence in World Models** — Introduces WROP, a dataset for training cognitive priors in video models. [HF Daily Papers](https://huggingface.co/papers/2609.28654)
- **How far behind Nvidia is Huawei?** — Analysis of Huawei’s AI compute constraints and scaling levers. [Epoch AI](https://epochai.substack.com/p/how-far-behind-nvidia-is-huawei)
- **The Chinese AI Infrastructure Boom** — Mapping of 1,000+ datacenter facilities across China. [SemiAnalysis](https://newsletter.semianalysis.com/p/the-chinese-ai-infrastructure-boom)
- **Boosting Agentic Coding with LLM Retries** — Lessons from Qwen3.8 27B on DeepSWE. [The Kaitchup](https://kaitchup.substack.com/p/boosting-agentic-coding-with-llm)
- **Audio8-ASR-Infinite** — A native streaming ASR model with unlimited-length transcription. [Hugging Face](https://huggingface.co/Edge0/Audio8-ASR-Infinite)
- **OpenRouter acquisition by Stripe** — Stripe bought OpenRouter for $7B. [Latent Space](https://www.latent.space/p/openrouter)
- **ClusterMAX 3.0** — Updated GPU cloud provider ratings for reliability and performance. [SemiAnalysis](https://newsletter.semianalysis.com/p/clustermax-30-the-industry-standard)
- **Fastest Qwen3.8 27B Quantization** — Benchmarking NVFP4, INT4, GSQ, and MTP on RTX Pro 6000. [The Kaitchup](https://kaitchup.substack.com/p/fastest-qwen38-27b-quantization-benchmarking)
- **Hemmingway-1** — A 27B model optimized for everyday writing tasks. [Hugging Face](https://huggingface.co/Altworld/Hemmingway-1)
- **WanPE** — A 397B prompt enhancement model for cinematic video generation. [HF Daily Papers](https://huggingface.co/papers/2609.30221)
- **Runway’s WorldPrompt** — Engineering real-time world models with persistent context. [Latent Space](https://www.latent.space/p/runway)
- **Note on 24th September 2026** — Simon Willison argues coding agents make software engineering harder. [Simon Willison](https://simonwillison.net/2026/Sep/24/harder/)
- **Bio-security is an AI Arms Race** — Radical Numerics uses multimodal AI for biological defense. [Latent Space](https://www.latent.space/p/bio-security-is-an-ai-arms-race-eric)
- **Gemini 3.8 TTS Playground** — Google releases new TTS models with voice cloning. [Simon Willison](https://simonwillison.net/2026/Sep/23/gemini-tts-playground/)
- **One Piece of Flock Camera Data** — Critical real-world harms of automated surveillance. [Jezebel](https://www.jezebel.com/flock-cameras-data-innocent-woman-arrested-lindsey-isaacs-palm-beach-florida-lawsuit-vehicular-homicide)
- **MiMo-V2.6-Distill-Qwen-9B** — A distilled agentic model for open research. [Hugging Face](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
- **strix** — Open-source AI penetration testing tool. [GitHub](https://github.com/usestrix/strix)
- **llm 0.36** — CLI tool update supporting GPT-6 Sol and Luna. [Simon Willison](https://simonwillison.net/2026/Sep/22/llm/)
- **How to keep enjoying programming in a world of LLMs** — Discussion on existential anxiety among engineers. [source](https://discourse.haskell.org/t/how-to-keep-enjoying-programming-in-a-world-of-llms/14705)
- **MiMo-V2.6-Flash-RL** — Efficiency-focused RL checkpoint. [Hugging Face](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
- **CLI-Anything** — Standardized interface for agent-native CLI tools. [GitHub](https://github.com/HKUDS/CLI-Anything)
- **ExplorationBench** — Benchmark for evaluating AI scientific exploration in verifiable alien worlds. [HF Daily Papers](https://huggingface.co/papers/2609.30199)
- **Parts-of-Speech as Emergent Categories in SAE Latent Space** — Insights into morpho-syntactic encoding in Sparse AutoEncoders. [HF Daily Papers](https://huggingface.co/papers/2609.29362)
- **Meta's Muse uses OpenAI model** — Reports suggest Meta’s Muse uses a model labeled muse-special. [source](https://mouse.dev/blog/muse-special/)
- **IterSynth** — Role-decoupled iterative synthesis for deep search agents. [HF Daily Papers](https://huggingface.co/papers/2609.29444)
- **A single function Jev-like wrapper for LLMs** — Technical pattern for standardizing LLM interfaces. [Allan Rbo](http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html)
- **Learning to Discover Interesting Mathematics** — Metric for interestingness in AI-generated theorems. [HF Daily Papers](https://huggingface.co/papers/2609.28603)
