PaperarXiv·29 Sep 2026·Inference efficiency
The loop: A language model natively manages its own context by updating a context file. The model's context-management instructions are then evolved through a skill-optimization loop, improving its accuracy and efficiency on subsequent tasks.
Language models are designed to natively manage their own context as a file, enabling intrinsic context management that can be optimized through skill evolution.
PaperarXiv·28 Sep 2026·Kernels and compilers
The loop: The KernelBraid agent optimizes unified RL kernels through an intermediate representation and verifies the improvements. These faster kernels are then used to accelerate the reinforcement learning training process.
An agentic framework optimizes unified RL kernels for rollout and policy update, preserving bitwise consistency and improving aggregate latency.
PaperarXiv·29 Sep 2026·Self-reward and self-play
The loop: A policy is temporarily trained ahead to create a stronger future teacher. This future teacher then provides dense token-level supervision to the restarted original policy, improving its performance.
A self-distillation method temporarily trains a model ahead to create a future teacher that provides better supervision for the original model.
PaperarXiv·29 Sep 2026·Tools and harness
The loop: Mutator agents rewrite complete agent harnesses using global search history and feedback. An orchestrator adapts the search strategy based on the performance of the discovered harnesses, improving the mutators' future harness designs.
A framework co-evolves agent harnesses and the mutator agents that design them, using hierarchical memory and orchestrated search adaptation.
PaperarXiv·30 Sep 2026·Training data
The loop: Self-evolving search agents generate their own training data, which often suffers from co-cheating where the proposer and solver agree on shared errors. This data is filtered using multi-sample verification and cross-fitting to reliably train the next generation of agents.
A study identifies and mitigates co-cheating in self-evolving search agents by verifying and partitioning self-generated training data.
PaperarXiv·29 Sep 2026·Self-modifying agents
The loop: An agent modifies its own instructions and tools using records of its previous self-improvement episodes. These modifications directly improve the agent's success rate on future downstream tasks without requiring external reward signals during the search.
A reward-free search procedure allows agents to modify their own instructions and tools using records of past self-improvement attempts.
PaperarXiv·29 Sep 2026·Self-reward and self-play
The loop: A policy generates its own feedback insights from failed attempts at difficult tasks. It then internalizes these insights through training, breaking through learning barriers and improving its success rate on future rollouts.
A reinforcement learning approach allows a policy to generate and internalize its own feedback insights from failed attempts to improve performance on difficult tasks.
PaperarXiv·1 Oct 2026·Training data
The loop: A model generates synthetic text which is then filtered using a non-parametric entropy rate estimator to preserve diversity. This filtered data is used to iteratively fine-tune the model, preventing model collapse across successive generations.
A non-parametric entropy rate estimator filters synthetic text to maintain diversity and prevent model collapse during iterative fine-tuning.