How can humans reliably oversee AI systems whose outputs, reasoning, or capabilities extend beyond what a human evaluator can directly verify? Scalable oversight is often presented as a catalogue of protocols—debate, critique, decomposition, weak supervision, panels, and recursive assistance. This book takes a different approach: it asks what an oversight procedure can actually justify.
Scalable Oversight of AI Systems develops a mathematical framework for analysing verification asymmetry, discriminating power, robust soundness, conditional completeness, capability gaps, distribution shift, and the limits of recursive supervision. It studies when decomposition preserves a defect, when process supervision helps or hurts, when weak supervision learns the wrong proxy, when many judges fail to add independent evidence, and how strategic assistance can reshape the distribution being evaluated.
The final chapters turn these results into safety-assurance objects: a claim must state its threat class, evaluation law, completeness class, review budget, premises, residual uncertainty, and expiry conditions. Throughout, formal guarantees are kept separate from measurements, heuristic models, engineering assumptions, and empirical findings.
Designed for advanced undergraduate and graduate courses, the volume combines rigorous derivations with worked system cases, CPU-only computational labs, structured exercises, research bridges, provenance tags, and reusable assurance templates. The result is a technical guide to a central question in advanced AI safety: not merely how to supervise systems beyond human expertise, but how to know what that supervision is entitled to claim.
"Sinopsis" puede pertenecer a otra edición de este libro.
Librería: California Books, Miami, FL, Estados Unidos de America
Condición: New. Nº de ref. del artículo: I-9798907070233
Cantidad disponible: Más de 20 disponibles