AI Agent AI工具 ALFWorld API-calling agent Adaptive Supervision Agent AgentFounder Agentic CPT Agentic Discovery Agentic RL Black-Box Distillation Claude Code Computer Use Continual Learning Continual Pre-training Credit Assignment DAPO DASH Deep Research EWC Echo Trap Forward KL GDPO GLM-4.5 GLM-5.3 GPT-6 Astra GRPD GRPO GSPO ICLR 2026 ICML 2026 Kimi K3 LLM LLM as world model LoRA Long-Horizon Long-Horizon Agent MOPD Meta-Learning MoE Multi-Turn RL Neural Computer OPD OPDVR OPSD On-Policy Distillation On-Policy-Distillation OpenAI PPO Post-Training Prompt Engineering RAG RAGEN REINFORCE RL RLHF RLVR SFT SMOPD SOPD Scaling Law Schmidhuber StarPO Step-Level Supervision Synaptic Intelligence Test-Time RL Test-Time Training Token Reweighting Tokenizer VLA VLM Video Model WAM World Model agentic RL agentic environment agentic flow agentic scaling agentic search attention continual learning continual reinforcement learning credit assignment curriculum curriculum learning data synthesis environment simulation environment synthesis flow matching function calling generative teaching graph hindsight instruction tuning model collapse model-based reinforcement learning multi-timescale planning process reward retriever training rollout self-evolution self-training slime survey synthetic data text world model token-level tool use tool-use agent value function value-free verified environment web agent world model 专家模型迭代 世界模型 代码生成 信息流 信用分配 具身智能 分层强化学习 参数正则化 可执行代码 后训练 基座模型 基模后训练 基础模型 多奖励优化 多教师蒸馏 多模态 奖励塑形 奖励门控 宏动作 对齐 工具使用 工程实践 开发工具 异步训练 强化学习 影响函数 技术报告 推理 推理模型 推理退化 数据合成 数据预处理 方差缩减 智能体 机器人 机器人学习 模型压缩 模型训练 混合推理 灾难性遗忘 物理推理 监督微调 知识蒸馏 稀疏奖励 组合泛化 细粒度奖励 网络安全 自动化 自蒸馏 计算范式 记忆训练 论文清单 论文综述 课程学习 软件架构 长上下文 长程任务 长程推理 预训练