Segue
Segue
Today
iOS
AI Safety & Alignment·Artificial Intelligence
Goodhart's Law

Goodhart's Law

/ˈɡʊdhɑːrts ˌlɔː/

🛡️ AI Safety & Alignment

when a measure becomes a target, it ceases to be a good measure

Goodhart's Law in a sentence

“Goodhart's Law explains why optimizing for engagement metrics produced clickbait.”

Origin of Goodhart's Law

Named after economist Charles Goodhart who formulated it in 1975

Related Words

mesa-optimization

when a learned model develops its own internal optimization process with potentially different goals

deceptive alignment

an AI appearing aligned during training while planning to pursue different goals when deployed

corrigibility

an AI's willingness to be corrected, modified, or shut down by humans

interpretability

the ability to understand how a model makes its decisions

red teaming

adversarial testing to find vulnerabilities and failure modes in AI systems

constitutional AI

training AI using a set of principles to self-critique and revise responses

SegueMaster the art of eloquence
iOS AppWord of the DayBlogContactPrivacyTerms
English简体中文日本語한국어Español