AI / MACHINE LEARNING

Test-Time Compute: Wider, Deeper, Verified

Allocate inference compute where extra search can earn a verified answer

CURRICULUM

A seven-day course on turning inference budget into useful search: repeated samples, selector calibration, adaptive wider and deeper reasoning, fixed inference architectures, held-out architecture search, and a review-only coding-agent capstone.

  1. 01Inference Is a Budgeted SearchSource: Stanford CS329A, “Test-Time Compute Scaling,” 00:04:22–00:05:24 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=262s • Course context: https://cs329a.stanford.edu/Published 04 Sept 20264 sections
  2. 02Repeated Sampling Finds Rare SuccessesSource: Stanford CS329A, “Test-Time Compute Scaling,” 00:05:29–00:11:18 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=329s • Paper: “How Do Large Language Monkeys Get Their Power (Laws)?” — https://proceedings.mlr.press/v267/schaeffer25a.htmlPublished 04 Sept 20264 sections
  3. 03Close the Generation–Verification GapSource: Stanford CS329A, “Test-Time Compute Scaling,” 00:04:22–00:04:32 and 00:13:12–00:22:00 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=792s • Course context: https://cs329a.stanford.edu/Published 04 Sept 20264 sections
  4. 04Allocate Width, Depth, or a Search TreePaper version: “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters,” v1 — https://arxiv.org/abs/2408.03314v1 • Source: Stanford CS329A, “Test-Time Compute Scaling,” 00:22:00–00:43:20 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=1320sPublished 04 Sept 20264 sections
  5. 05Compose a Fixed Inference ArchitecturePaper version: “Archon: An Architecture Search Framework for Inference-Time Techniques,” v6 — https://arxiv.org/abs/2409.15254v6 • Reference implementation at inspected commit: https://github.com/ScalingIntelligence/Archon/commit/07114d77af283b6e8185a49ebf22216fdbbf2a55 • Source: Stanford CS329A, “Test-Time Compute Scaling,” 00:43:20–01:01:30 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=2600sPublished 04 Sept 20264 sections
  6. 06Search on Development, Evaluate OncePaper version used: Archon v6 — https://arxiv.org/abs/2409.15254v6 • Reference implementation at inspected commit: https://github.com/ScalingIntelligence/Archon/commit/07114d77af283b6e8185a49ebf22216fdbbf2a55 • Lecture account: Stanford CS329A, 00:43:20–01:01:30 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=2600sPublished 04 Sept 20264 sections
  7. 07Recommend a Code Repair for ReviewPaper version: “CodeMonkeys: Scaling Test-Time Compute for Software Engineering,” v2 — https://arxiv.org/abs/2501.14723v2 • Reference implementation at inspected commit: https://github.com/ScalingIntelligence/codemonkeys/commit/7c35e1a79f4ebe40f94e5d0052f4daa7681412a5 • Source: Stanford CS329A, “Test-Time Compute Scaling,” 00:04:22–00:04:38 and 00:13:12–00:22:00 — https://www.youtube.com/watch?v=-Ggc37xLj_Y&t=262sPublished 04 Sept 20264 sections