Tech Blog

AI Is Improving AI Faster Than We Can Check

9 min read

AI is already helping to build the next generation of AI, with people still in the loop. What limits its speed is verification: agents produce work far faster than people and tests can check it. Five reasons verification is the...

Where AI Self-Improvement (RSI) Stands: 8 Takeaways / AI 自我改进(RSI)走到哪一步了:8 条 take away

17 min read

On October 3, NICE hosted an online workshop on the mechanisms, evidence and limits of AI self-improvement, with four researchers who work on RSI. The main approaches, the current bottlenecks, how to evaluate it, and what it means for academia...

MatrAIx: The Hardest Part of AI Evaluation / MatrAIx:AI评测中最难的部分

16 min read

Behind 8.3 billion simulated users, a startup is targeting the part of AI evaluation that benchmarks never measured. An interview with Xiaomin Li and Yuexing Hao, the two founders of MatrAIx: how the personas get built and turned into agents,...

英伟达为什么要保卫开放权重,以及它将如何重塑AI算力市场 / Why NVIDIA Is Defending Open-Weight Models—and How They Could Reshape the AI Compute Market

17 min read

Open weights are not just a model-release choice; they reshape who buys compute, who controls inference demand, and how much bargaining power enterprises have against closed APIs.

The Evolution of Agents: From Context Engineering to Long-running Harnesses / Agent 从 Context Engineering 到 Long-running Harness 的演变过程

28 min read

From LLM + tool use to context engineering, and then to long-running agent harnesses: agent capability is becoming a system property composed of the model, harness, context, tools, evals, sandbox, and state management.