2026
-
CS336: Lecture 2 - PyTorch and resource accounting
Lecture 2 is about making training cost concrete: tensors, dtypes, memory, FLOPs, autograd, optimizers, data loading, checkpoints, and mixed precision all have resource prices.
-
AMP: automatic mixed precision as a dispatch policy
TLDR: AMP is not "turn the model into half precision." It is a runtime policy that runs safe, high-throughput ops in lower precision while protecting numerically sensitive paths.
-
Autocurricula and Multi-Agent Innovation: 社会互动如何生成新问题
TLDR: Multi-agent intelligence should study how cooperation, competition, specialization, and shared discoveries create abilities that isolated agents would miss.
-
Social Dilemmas: 三个经典社会困境
TLDR: Social dilemmas show why individually rational actions can damage group outcomes, and why cooperation depends on payoffs, repetition, reputation, and norms.
-
A Social Path to Human-Like AI: 社会互动如何生成新数据
TLDR: Human-like AI may require populations of agents learning through social interaction, where cooperation and competition generate skills beyond single-agent training.
-
Talk with Shunyu Yao: feedback is the center of AI research
TLDR: The conversation is useful because it frames AI research as system-driven experimental work: define verifiable problems, build feedback loops, debug carefully, and choose directions where scaling paths are still being shaped.
-
Anthropic Blogs: harness engineering and context engineering
The shared lesson across these Anthropic engineering posts is that long agent tasks fail at the runtime layer: context, evaluation, sandboxing, permissions, handoff, and feedback have to be engineered.
-
Building a C compiler with agent teams
The C compiler experiment worked because the project had the right substrate for agents: a modular architecture, objective tests, Git as shared memory, task locks, readable logs, and oracles that turned one giant goal into many local failures.