← All paper explainers  ·  ravikant.dev

Intrinsic (Short-Term) Reward in RL — a ranked field map

An exhaustive, ranked reading list of the intrinsic reward literature: the short-term signal an agent gives itself — curiosity, novelty, surprise, empowerment — layered on top of the sparse long-term reward from the world.

total reward  =  long-term environment reward  (sparse, after many actions)  +  short-term intrinsic reward  (dense, self-generated)

How to read this

This page maps the intrinsic side. Everything is ranked as a single learning path — read top-to-bottom and the field assembles itself: foundations first, then the core curiosity/novelty methods, the major branches (information gain, empowerment, state-entropy, episodic memory, learning progress, homeostasis), then specialized work, theory/surveys, and the LLM-era bridge back to the papers in this collection. Filter by sub-branch below. Summaries get to the mechanism and stop.

🔥 Want the concepts, not the list? The approaches & the math that separates them ↗ a synthesis organized by method, with each reward formula explained simply. This page is the exhaustive bibliography; that one is how to think about the subject.

Is this map useful?

Tell me what's missing, mis-ranked, or worth a full deep-dive. Anonymous.

Useful field map?
Missing papers, wrong ranking, or which one to deep-dive next?
Thanks — logged.