## The big picture

**Aleph Alpha Kolibri**
Aleph Alpha has released Kolibri, a sovereign open-weight large language model developed in Germany.
[Blog](https://tej.as/blog/aleph-alpha-kolibri)
**OpenAI Safety Resignation**
A high-profile safety team member has resigned from OpenAI, citing a broken internal culture. The departure highlights ongoing tensions regarding safety priorities and corporate direction within the leading AI lab.
[source](https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?gift=v5U_UzUTothfWXsPxtvNVAh7esWToMRD6XnbXmc5WgA)

## Architectural breakthroughs

**LESSER: Output-Layer Gradients for Data Selection**
The authors propose LESSER, a post-training data selection method that approximates full-parameter gradient alignment using only output-layer gradients. By requiring only a forward pass instead of a full backward pass, it reduces feature-extraction FLOP costs by 9.7x for SFT while maintaining performance comparable to expensive full-gradient methods.
[arXiv](https://arxiv.org/abs/2610.03702v1)

**Homa: The end of TCP for AI clusters**
Homa is presented as the end of TCP for AI clusters.
[source](https://www.youtube.com/watch?v=eZ8WWZzoaR0)
**Pivot-SD: Self-Distillation for Masked Diffusion**
Pivot-SD addresses credit assignment in masked diffusion language models by supervising only high-impact token commitments. It uses an information-gain metric to identify pivots that significantly reduce uncertainty over remaining masked positions, enabling efficient offline self-distillation.
[arXiv](https://arxiv.org/abs/2610.03665v1)

**Local Support Learning**
This framework tackles catastrophic forgetting by treating it as a geometric problem in the input space of weight matrices. It introduces a gating function that restricts new updates to the local support of the new training distribution, preserving prior capabilities without requiring access to previous data.
[HF Daily Papers](https://huggingface.co/papers/2610.02126)

## New open-weight model releases

## Hardware & optimization

**MRVQ: Elastic Vector Search**
Matryoshka Residual Vector Quantization (MRVQ) allows a single index artifact to serve varying dimension and bit-rate requirements. By truncating residual stages or embedding coordinates, it reduces memory usage by up to 22x compared to separately trained indices, enabling dynamic adaptation to latency and quality budgets.
[arXiv](https://arxiv.org/abs/2610.03651v1)

**AutoTarget: Solver-Aware Caching**
AutoTarget optimizes caching in Diffusion Transformers by selecting the most appropriate tensor to reuse based on the specific solver and reuse schedule. It measures error from candidate tensors without cache reuse to minimize degradation in few-step generation, improving efficiency without compromising quality.
[arXiv](https://arxiv.org/abs/2610.03577v1)

## Also this week

- **Fold2Reason**: Post-training on protein folding data improves general spatial reasoning benchmarks. [HF Daily Papers](https://huggingface.co/papers/2609.38879)
- **HyperBrowseComp**: A multilingual, multimodal benchmark for web-browsing agents requiring complex evidence retrieval. [arXiv](https://arxiv.org/abs/2610.03574v1)
- **MotorMind**: Scaffolds general VLMs for zero-shot robot manipulation without external action experts. [HF Daily Papers](https://huggingface.co/papers/2609.38078)
- **4DCodeBench**: Benchmarks agents on reconstructing dynamic 3D scenes from video as executable graphics code. [arXiv](https://arxiv.org/abs/2610.03715v1)
- **Stratified Retention**: Proposes separating invariant knowledge from non-stationary facts in continual world models. [arXiv](https://arxiv.org/abs/2610.03713v1)
- **NAVA-WAM**: Learns action priors directly from observation-only videos for world action models. [HF Daily Papers](https://huggingface.co/papers/2610.03391)
- **EyeRobot 2.0**: Uses active gaze and foveal processing for precise bimanual manipulation with a single stereo camera. [arXiv](https://arxiv.org/abs/2610.03710v1)
- **SciUtopia**: Simulates academic research ecosystems using closed-loop LLM agents. [HF Daily Papers](https://huggingface.co/papers/2610.01257)
- **Queen**: A 4B-parameter chess-language model that explains moves at Grandmaster level. [arXiv](https://arxiv.org/abs/2610.03695v1)
- **Spatial Memory Intelligence**: Manages long-term spatial context in world models using MLLM reasoning. [HF Daily Papers](https://huggingface.co/papers/2610.02521)
- **FrugalEvo**: Cost-aware LLM-guided program evolution optimizing for gain per unit cost. [arXiv](https://arxiv.org/abs/2610.03675v1)
- **ProAR**: Introduces prospective reasoning to autoregressive video models for goal-oriented generation. [HF Daily Papers](https://huggingface.co/papers/2610.03664)
- **Planning to Learn**: Theoretical analysis showing cross-entropy outperforms exact policy gradients in classification. [arXiv](https://arxiv.org/abs/2610.03667v1)
- **Efficient CoT**: Investigates trade-offs between reasoning efficiency and faithfulness in Chain-of-Thought. [HF Daily Papers](https://huggingface.co/papers/2610.03509)
- **Forecasting from Counterfactuals**: Evaluates Sim2Real transfer for forecasting models under new decision policies. [arXiv](https://arxiv.org/abs/2610.03662v1)
- **VeriHarness**: Scales agentic verification for long-horizon tasks using disagreement and consensus checking. [HF Daily Papers](https://huggingface.co/papers/2610.00972)
- **Looped MoE**: Addresses obstacles to scaling looped Mixture-of-Experts beyond two iterations. [HF Daily Papers](https://huggingface.co/papers/2610.01153)
- **Success Conditioning**: Proves convergence rates for success conditioning in RL policies. [arXiv](https://arxiv.org/abs/2610.03642v1)
- **Triadic Linear Attention**: Extends linear attention to 3D tensor states for long-context modeling. [HF Daily Papers](https://huggingface.co/papers/2609.36529)
- **IDRF**: Reward fine-tuning for masked discrete diffusion models using inverse-distillation regularization. [arXiv](https://arxiv.org/abs/2610.03641v1)
- **ProWAM**: Progressive world action models predicting sparse visual sub-goals for long-horizon control. [HF Daily Papers](https://huggingface.co/papers/2610.02508)
- **LoGo**: Local-global rewards for consistent long-horizon video generation. [arXiv](https://arxiv.org/abs/2610.03636v1)
- **DepGPO**: Dependency-aware policy optimization for terminal agents using command graphs. [arXiv](https://arxiv.org/abs/2610.03634v1)
- **HelixWorld**: Real-time interactive audio-visual world model with spatial stereo sound. [HF Daily Papers](https://huggingface.co/papers/2609.38123)
- **World Embedding Benchmark**: Evaluates physical fidelity in video representations across mechanics and optics. [arXiv](https://arxiv.org/abs/2610.03632v1)
- **Depth as Time**: Empirical observation that denoising trajectories unfold across network depth in one-step models. [arXiv](https://arxiv.org/abs/2610.03626v1)
- **Dream4ACT**: Shared visual action interface for multi-embodiment video-action modeling. [HF Daily Papers](https://huggingface.co/papers/2609.40153)
- **FALCON**: Framework for generating realistic, ambiguity-aware synthetic NL2SQL data. [arXiv](https://arxiv.org/abs/2610.03625v1)
- **UniIntervene++**: Adaptive intervention agent for efficient real-world robot learning. [arXiv](https://arxiv.org/abs/2610.03620v1)
- **Wayfarer**: Discovers temporal abstractions (options) in high-dimensional RL domains. [arXiv](https://arxiv.org/abs/2610.03604v1)
- **PI-SME**: Path integral surrogate for multi-step gradient inversion in federated learning. [arXiv](https://arxiv.org/abs/2610.03597v1)
- **TPRS**: Measures threat-preserving representation sensitivity in agent security benchmarks. [arXiv](https://arxiv.org/abs/2610.03585v1)
- **SCM**: AI search for photos and video frames on macOS. [GitHub](https://github.com/allenv0/SCM)
- **Experiential**: Open-source gateway for model routing and cost optimization. [GitHub](https://github.com/experientiallabs/experiential)
