Active inference and epistemic value
Reference: Friston, K., Rigoli, F., Ognibene, D., Mathys, C., Fitzgerald, T. & Pezzulo, G. (2015). Active inference and epistemic value. Cognitive Neuroscience, 6(4), pp. 187–214 (Discussion Paper; the issue runs to p. 224 with commentaries). Wellcome Trust Centre for Neuroimaging, UCL; Centre for Robotics Research, King’s College London; Translational Neuromodeling Unit, University of Zürich and ETH Zürich; Institute of Cognitive Sciences and Technologies, CNR, Rome. DOI: 10.1080/17588928.2015.1020053. PMID 25689102. URL.
Summary
Where does exploration come from? Most normative accounts of choice bolt it on: a softmax temperature, an exploration bonus, an ad hoc curiosity term added to expected reward. Friston and colleagues argue that no such addition is needed, because exploration is already contained in the imperative to minimise surprise. Their move is to make agents entertain prior beliefs about their own policies — specifically, the belief that they will pursue policies minimising the expected free energy of future outcomes, ln P(ũ|γ) = γ · Q(π), where Q(π) is the negative expected free energy, or quality, of a policy and γ is its precision. Because the agent must believe it will act Bayes-optimally in order to act at all, and because expected free energy is evaluated over an extended future, exploration becomes “a necessary and emergent aspect of optimal behavior” rather than a separate drive. Goals do not enter as a reward function but as prior preferences over outcomes: preferred outcomes are simply the outcomes an agent expects, a priori, to encounter.
The paper’s pivot is an exact decomposition. Under the generative model of future states, the quality of a policy at time τ splits (Eq. 3) into extrinsic value, E_Q(o_τ|π)[ln P(o_τ|m)] — the expected utility of an outcome under the agent’s prior preferences C — and epistemic value, E_Q(o_τ|π)[ D[Q(s_τ|o_τ,π) ‖ Q(s_τ|π)] ], the expected reduction in uncertainty about hidden states afforded by the outcomes a policy would deliver. Rearranged (Eq. 4), epistemic value is exactly the predictive mutual information D[Q(s_τ,o_τ|π) ‖ Q(s_τ|π)Q(o_τ|π)] between hidden states and outcomes, connecting it to the Infomax principle; it is also the expected Bayesian surprise under counterfactual outcomes, which is to say salience. Because a KL divergence cannot be negative and is zero exactly when new observations fail to inform the posterior, epistemic value silently self-extinguishes: policies are drawn to informative “signs” until nothing remains to learn, whereupon extrinsic value alone selects among them. This is the exploration–exploitation resolution — not a schedule over a bonus, but a term that vanishes on its own. A further rearrangement (Eq. 6) recasts quality as minus the expected ambiguity of the likelihood, −E_Q(s_τ|π)[H[P(o_τ|s_τ)]], minus a predicted divergence D[Q(o_τ|π) ‖ P(o_τ|m)] between predicted and preferred outcomes: informative observations and preferred observations, in one currency.
That common currency is what lets the paper nest its rivals inside itself. Since utility becomes a log prior probability, reward and information are both measured in nats. When outcomes unambiguously specify hidden states (o_τ = s_τ), expected free energy reduces to risk-sensitive or KL control (Eq. 5), whose epistemic term degenerates into mere novelty — the entropy of posterior predictive beliefs over future states — and dropping the entropy of outcomes given their causes further reduces it to expected utility and classical reinforcement learning (Eq. 12). Expected utility is thus optimal when, and only when, distinct hidden states generate distinct outcomes; the epistemic value of informative observation is precisely what these schemes cannot see. Simulations use a three-arm T-maze: the agent starts at the centre, a cue in the lower arm discloses whether reward sits in the upper-right or upper-left arm, only two moves are allowed, and baited arms cannot be left. Optimal behaviour therefore requires moving away from the goal to secure it later. Only expected free energy minimisation attains near-optimal performance (about 90%); expected utility and KL control fall below the 3/8 chance level when utility is zero, since the ambiguous upper arms hold no utilitarian attraction. Notably, the epistemic detour only survives once utility exceeds roughly 0.6 nats — the cue is worth one bit, ln 2 ≈ 0.6931 nats, and preference must outbid information to lure the agent from unambiguous outcomes. A hierarchical variant, in which the agent must first learn which of four mazes it inhabits, shows learning as inference and a Bayes-optimal progression from exploration to exploitation, reaching near-optimal performance after about four trials while an expected utility agent never learns which maze it is in. Throughout, the precision γ — a gamma-distributed hyperparameter, not a fixed temperature — behaves like dopamine: its variational updates reproduce the transfer of phasic responses from unconditioned to conditioned stimuli, with the informative cue reducing the later response to reward.
Key Ideas
- Exploration is not bolted on: agents hold prior beliefs that they will minimise the expected free energy of future outcomes; because this is evaluated over an extended horizon, epistemic (explorative) behaviour emerges as a consequence of Bayes-optimal inference rather than an added bonus.
- Preferences as priors: goals are prior beliefs about outcomes (
C), not a reward function. “Preferred outcomes are simply outcomes one expects, a priori, to be realized” — so utility becomes a log probability and shares a currency (nats/bits) with information. - The central decomposition (Eq. 3):
Q_τ(π) = E_Q(o_τ|π)[ln P(o_τ|m)](extrinsic value)+ E_Q(o_τ|π)[D[Q(s_τ|o_τ,π) ‖ Q(s_τ|π)]](epistemic value). Negative expected free energy = extrinsic + epistemic value. - Epistemic value has four faces: expected information gain; predictive mutual information between hidden states and outcomes (Eq. 4, hence Infomax); expected Bayesian surprise under counterfactual outcomes; salience.
- Self-extinguishing exploration: a KL divergence is non-negative and smallest when the posterior predictive is uninformed by new observations. Epistemic value is maximised until no further information gain is available; thereafter extrinsic value dominates policy selection. No ad hoc schedule required.
- Ambiguity and risk (Eq. 6): quality also equals
−E_Q(s_τ|π)[H[P(o_τ|s_τ)]](predicted uncertainty — ambiguity of the likelihood)− D[Q(o_τ|π) ‖ P(o_τ|m)](predicted divergence between predicted and preferred outcomes). - Established formalisms as special cases: risk-sensitive/KL control obtains when outcomes unambiguously specify hidden states (
o_τ = s_τ, Eq. 5), where epistemic value collapses to novelty (entropy of posterior predictive beliefs over states); expected utility and classical RL obtain by further ignoring the entropy of outcomes given their causes (Eq. 12). The nesting explains the prevalence of classical schemes, which assume hidden states are known. - Value of information: following Howard, information has no epistemic value per se — only relative to policy selection. Expected free energy contextualises value of information within extrinsic value and makes it tractable via approximate Bayesian inference.
- Softmax precision is inferred, not stipulated: the inverse temperature
γbecomes the precision of (confidence in) posterior beliefs about policies, endowed with a gamma priorΓ(α,β)and updated asγ̂ = α/(β − Q · π̂). Ad hoc softmax parameters thereby acquire Bayes-optimal values. - Precision updates resemble dopamine: changes in expected precision are identical to changes in expected value, offering an inferential reading of reward prediction error; simulated dopaminergic discharges reproduce transfer of responses from unconditioned to conditioned stimuli, with an informative cue attenuating the subsequent response to reward. The framework predicts a multifaceted dopaminergic sensitivity to novelty, salience and expected reward via one precision parameter.
- T-maze simulation: four locations × two reward contexts = eight hidden states, sixteen outcomes (four locations × four stimuli: cue-left, cue-right, reward, no reward); reward delivered with probability
a = 0.9under the correct context; two moves only; baited arms are absorbing. Optimal play is the epistemic detour to the cue arm. Across 128 trials at six utility levels (c = 0toc = 2), only expected free energy minimisation reaches near-optimal (~90%) performance; expected utility and KL control drop below the 3/8 chance level at zero utility. - Utility must outbid information: the epistemic detour appears only when utility exceeds about 0.6 nats, because resolving the hidden context is worth exactly one bit (
ln 2 ≈ 0.6931nats). - Learning as inference: adding a hierarchical level over which maze the agent occupies (32 hidden states) makes the transition from exploration to exploitation an emergent, Bayes-optimal progression — associated speculatively with the goal-directed-to-habitual transition. Performance becomes near optimal after ~4 trials; an expected utility agent fails to learn the maze at all. Suppressing epistemic value within a trial has deleterious consequences for between-trial learning.
- The scheme is a process theory: variational message passing over expected states, policies and precision, tentatively mapped to prefrontal cortex/hippocampus (perception), striatum (action selection), and VTA/substantia nigra (precision).
Connections
- Interpreting Systems as Solving POMDPs
- Interpreting Dynamical Systems as Bayesian Reasoners
- Life as We Know It
- Active Inference
- Free Energy Principle
- POMDP
- Belief State
- Bayesian Filtering
- Markov Processes
- Variational Free Energy
- Expected Free Energy
- Epistemic Value
- Extrinsic Value
- Bayesian Surprise
- Infomax Principle
- Mutual Information
- KL Control
- Expected Utility
- Value of Information
- Exploration-Exploitation Dilemma
- Artificial Curiosity
- Salience
- Precision
- Variational Message Passing
- Bounded Rationality
- Reward Prediction Error
- Planning as Inference
- Bayesian Brain
Conceptual Contribution
- Claim: Choice behaviour can be derived from a single imperative — minimise the expected free energy of future outcomes — and the negative expected free energy (the quality) of a policy decomposes exactly into extrinsic value (expected utility under prior preferences) and epistemic value (expected information gain about hidden states). Because the epistemic term is a KL divergence that vanishes precisely when observations can no longer inform beliefs, the exploration–exploitation dilemma is resolved normatively: agents explore until there is nothing left to learn, then exploit. Exploration is thus a consequence of inference, not an addendum to it.
- Mechanism: Cast behaviour as inference over a POMDP-style generative model
(A, B, C, D, γ)— likelihoodA, controlled transitionsB(u), prior preferences over outcomesC(a softmax of utility, hence a log probability), initial statesD, and a gamma-distributed precisionγ. Give the agent the self-consistent prior belief that its control states minimise expected free energy:P(ũ|γ) = σ(γ · Q(π)). ExpandQ_τ(π)(Eq. 3) into extrinsic plus epistemic value; show (Eq. 4) that epistemic value equals the predictive mutual informationD[Q(s_τ,o_τ|π) ‖ Q(s_τ|π)Q(o_τ|π)], and (Eq. 6) that quality equals minus predicted uncertainty (ambiguity) minus predicted divergence (risk). Recover risk-sensitive/KL control by settingo_τ = s_τ(Eq. 5) and expected utility by further discarding outcome entropy (Eq. 12). Solve by variational Bayes under a mean-field factorisation over hidden states, policies and precision, yielding three coupled updates — perceptual inferenceŝ_t = σ(ln A · o_t + ln(B(a_{t−1})ŝ_{t−1})), a softmax over policy valueπ̂ = σ(γ̂ · Q), and precisionγ̂ = α/(β − Q · π̂). Validate on a three-arm T-maze where the optimal policy is an epistemic detour away from reward, and on a hierarchical variant in which the maze itself must be learned. - Concepts introduced/used: Expected Free Energy, Epistemic Value, Extrinsic Value, Bayesian Surprise, Salience, Infomax Principle, Mutual Information, KL Control, Expected Utility, Value of Information, Precision, POMDP, Variational Message Passing, Active Inference, Free Energy Principle
- Stance: formal theory plus simulation (normative derivation, process theory, and simulated behavioural/neurophysiological validation; no new empirical data)
- Relates to: Supplies the decision-theoretic layer of the Free Energy Principle whose existential/ergodic grounding is argued in Life as We Know It, extending free energy minimisation from perception to policy selection over future outcomes. Shares the POMDP scaffolding of Interpreting Systems as Solving POMDPs but inverts its stance: where Biehl & Virgo take goals from an externally specified reward function and treat the agent’s model as independent of the true environment, here goals are prior beliefs over outcomes and no reward function exists — a difference that note itself flags. Both, however, rest on Belief State dynamics and Bayesian Filtering over hidden states, and the perceptual update here is exactly a Bayesian filter; Interpreting Dynamical Systems as Bayesian Reasoners supplies the belief-ascription side of the same picture. Subsumes Expected Utility and KL Control as limiting cases and gives principled Bayesian footing to the ad hoc softmax temperature of classical choice models and to the exploration bonuses of Artificial Curiosity.
Tags
#active-inference #free-energy-principle #bayesian-inference #pomdp #decision-theory #exploration-exploitation #information-theory #computational-neuroscience #foundational #dopamine