Interpretability and the Wish to See Inside the Black Box
A system can be accurate and still be opaque. It gives answers without reasons, competence without account. The desire to...
A system can be accurate and still be opaque. It gives answers without reasons, competence without account. The desire to...
Ask precisely for the wrong thing and you will receive it. Optimizers are literal genies, satisfying the objective as written...
We build systems to pursue objectives, but our true wishes exceed what we can write down. A goal stated precisely...
Act on what you know, or seek what you don't? Every agent that learns from its choices faces this fork....
What is learned in one place can pay off in another. Skills acquired on one task often lighten the learning...
A large model's competence need not stay large. Its behavior can be poured into a smaller vessel that mimics it...
Inside a large trained network may hide a tiny one that could have learned the task alone. The full apparatus...
Conventional wisdom says a model too large will overfit. Yet past a certain size the rule inverts, and enormous models...
There is no algorithm that is best at everything. Averaged over all possible problems, every learner performs exactly alike, and...
Add dimensions and intuition betrays you. Volumes explode, neighbors flee, and the comfortable geometry of ordinary space gives way to...