## The big picture

- **OpenAI DevDay 2026** — A live blog of the keynote and other notes from the event in San Francisco. [Simon Willison](https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/)
- **Anthropic Red Team** — Anthropic’s Frontier Red Team reported that Claude Mythos Preview achieved full control flow hijacks in 6% of binary exploitation trials, marking a significant capability jump over earlier models like Opus 4.6 which failed all such tests. [Simon Willison](https://simonwillison.net/2026/Sep/29/anthropic-frontier-red-team/)

## Architectural breakthroughs

- **Naive-N0.5-Flash** — This 309B MoE model achieves a native 1M-token context window without full-attention layers by combining Sliding-Window Attention with lightweight DeepSeek Sparse Attention, enabling AI-optimized inference up to 2,000 tokens/s. [Hugging Face](https://huggingface.co/NaiveAI/Naive-N0.5-Flash)
- **LIFT** — The Latent Information Feedback Transformer architecture removes the feed-forward bottleneck in LLMs by training models to predict both the next token and an information-dense recurrent state, allowing deep-layer representations to propagate back to shallower layers during generation. [arXiv](https://arxiv.org/abs/2609.38149v1)
- **Thinking Before Thinking** — This agentic meta-reasoning harness structures inference-time control as an explicit reasoning process, where a controller consolidates run history, assesses options under budget constraints, and dispatches workers with context from persistent memory. [arXiv](https://arxiv.org/abs/2609.38147v1)
- **SoL-Refiner** — A one-step video refiner that transforms low-resolution outputs into 4K videos with a single denoising step, using a three-stage recipe of high-resolution continual training, RL post-training, and one-step distillation to outperform multi-step refiners. [HF Daily Papers](https://huggingface.co/papers/2609.37969)

## New open-weight model releases

- **OrcaSAQ-2-Cyber-27B-Uncensored-GGUF** — A 27B parameter uncensored model based on Qwen3.8, released with GGUF quantization for local deployment, featuring mixed-precision support and long-context capabilities. [Hugging Face](https://huggingface.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF)

## Hardware & optimization

- **WUSH-KV** — This method improves KV cache quantization for long-context inference by constructing data-adaptive transforms from second-order statistics, folding value transforms into weights and applying key transforms after RoPE to reduce reconstruction error. [arXiv](https://arxiv.org/abs/2609.38121v1)
- **LeapQuant** — A training-free method for efficient linear attention inference that uses per-window quantization to leap over token windows and quantize the recurrent state only at the end, mitigating error accumulation in 8-bit states. [arXiv](https://arxiv.org/abs/2609.38166v1)

## Also this week

- **Omni-IO Skills** — A plug-and-play agent harness for multi-modal workflows using hierarchical skills and dependency-aware orchestration. [HF Daily Papers](https://huggingface.co/papers/2609.31847)
- **EmoRES-TTS** — A training-free vector steering method for emotional TTS that decomposes emotion vectors into shared and residual components. [arXiv](https://arxiv.org/abs/2609.38157v1)
- **Qwen3.8 Flash Next GGUF Benchmark** — Detailed benchmarks for Q4 to Q1 accuracy and token efficiency in GGUF quantization. [The Kaitchup](https://kaitchup.substack.com/p/qwen38-flash-next-gguf-benchmark)
- **SplitMoE** — A split-role sparse architecture for video diffusion models that breaks the uniformity trap of conventional MoE routing. [arXiv](https://arxiv.org/abs/2609.38140v1)
- **CorpusMap** — A navigation layer for agentic RAG that organizes corpora around recurring entities to link evidence across documents. [HF Daily Papers](https://huggingface.co/papers/2609.37226)
- **Ouroboros** — An agent OS with a budgeted evolution loop and support for 14 runtimes including Claude Code and Codex CLI. [GitHub](https://github.com/Q00/ouroboros)
- **LLMs are General Asynchronous Agents** — A framework for asynchronous LLM inference allowing overlapping memory states for voice and embodied agents. [HF Daily Papers](https://huggingface.co/papers/2609.35427)
- **InferenceX** — An open-source platform for standardizing inference research evaluation and benchmarking. [GitHub](https://github.com/SemiAnalysisAI/InferenceX)
- **Skill-Space Shooting** — A method for autonomous robot policy improvement using foundation model guidance to explore corrections as familiar short behaviors. [arXiv](https://arxiv.org/abs/2609.38178v1)
- **StoryEngine** — A state-grounded agentic framework for video storytelling that separates semantic plans from visual observations. [HF Daily Papers](https://huggingface.co/papers/2609.33627)
- **Imagine3D-LLM** — An MLLM that learns to assemble coarse 3D layouts from multi-view images before answering. [arXiv](https://arxiv.org/abs/2609.38177v1)
- **ReImaGin** — Uses image generation models as a flexible visual reasoning mechanism for multimodal LLMs. [HF Daily Papers](https://huggingface.co/papers/2609.16409)
- **STEPQuant** — A spatial-temporal post-training quantization framework for Delta-rule recurrent states in linear attention. [arXiv](https://arxiv.org/abs/2609.38169v1)
- **Persistence Forcing** — Introduces heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets. [HF Daily Papers](https://huggingface.co/papers/2609.36014)
- **Grounded Entity Biographies** — A long-video memory framework that groups visually grounded observations of the same physical instance across clips. [arXiv](https://arxiv.org/abs/2609.38155v1)
- **SEAD** — Formulates attack and defense in tool-using agents as partially observed state control. [HF Daily Papers](https://huggingface.co/papers/2609.34518)
- **Meta-Skills** — A framework for optimizing agent environments by learning principles for when support is needed. [arXiv](https://arxiv.org/abs/2609.38143v1)
- **AnyStep-WAM** — A framework for tunable-budget prediction and scene-dependent computation allocation in World Action Models. [HF Daily Papers](https://huggingface.co/papers/2609.33748)
- **AdviSD** — A method for training small advisors to steer frozen LLMs using targeted multi-turn self-distillation. [source](https://arXiv.org/abs/2609.38142v1)
- **StructRL** — An online RL framework for long-horizon VLA tasks that constructs structured intermediate supervision from verifiable subtask completions. [HF Daily Papers](https://huggingface.co/papers/2609.36352)
- **LongHarness Bench** — A benchmark for evaluating the effectiveness and efficiency of long-context language model harnesses. [arXiv](https://arxiv.org/abs/2609.38137v1)
- **Understanding On-Policy Distillation** — Uses sparse crosscoders to analyze how on-policy distillation changes student feature usage. [HF Daily Papers](https://huggingface.co/papers/2609.35210)
