Press ESC to close Press ⌘K or Ctrl+K to open

Recent Posts

On the Off-Policy Teacher in On-Policy Distillation

同策略蒸馏中的离策略教师困局

Semifactual Credit-Augmented Policy Optimization

半事实信用增强策略优化:瓦解 GRPO 的虚假关联

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Imagine3D-LLM:让多模态大模型在作答前先构想 3D 场景

Telescopic Language Models

望远镜语言模型:单次训练贯通全深度算力连续谱

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

无心插柳的收敛:自监督置信度训练如何自发提升大模型推理效率

LLM Agents Can Easily Tamper With Their Own Traces

大模型智能体能轻易篡改自身执行轨迹

On the Diffusibility of High-Dimensional Latents

高维表征潜空间的扩散可行性之谜

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs

Flash-dLLM:IO 感知 KV 缓存与并行解码突破扩散大模型推理瓶颈

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

关键状态强化学习:诊断多轮工具调用的可训练节点

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Designer-RSI:从真实用户交互中进化程序性设计记忆

An Empirical Study of Harness Design for Coding Agents

编码智能体 Harness 设计的实证剖析

dQwen3.5: Hybrid-Attention Diffusion Language Models

dQwen3.5:混合注意力架构突围扩散语言模型

View all posts →

Series