Learning What to Recall:
Adaptive Multi-Cue Episodic Memory for World Models

Learning which memories to recall, and which cues to trust, from the future the world model is trained to predict.

Beomsu Kim1,2, Chieh-Hsin Lai2, Bac Nguyen2, Amir Bar3, Jong Chul Ye1,†, Yuki Mitsufuji2,†

1Korea Advanced Institute of Science & Technology  ·  2Sony Group Corporation  ·  3Imperial College London
†Corresponding authors

Under review

Abstract

World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question:

Which memories are useful for the current prediction, and which retrieval cues should be trusted to find them?

Fixed criteria based on recency, pose overlap, or visual similarity are unreliable across environments and queries. Future-Aware Recall (FAR) instead learns recall: during training, the realized future reveals which memories actually helped the prediction, and that signal trains a retriever that stays future-blind at inference and learns which cues to trust for each query.

Across three complementary settings, FAR outperforms hand-designed recall even with the same retrieval cues, automatically adapts which available cues to trust, and recalls the right history as the world changes. Together, these results establish FAR as a flexible, principled approach to episodic memory access in world models.

Method

Future-Aware Recall

Overview of FAR: cue-specific scorers over the interaction history, query-dependent cue fusion, Top-K recall, and world-model prediction, with predictive utility from the observed future supervising the scorers during training only.
During training, the realized future supervises the scorers through predictive utility; at inference, recall is future-blind.

1Predictive utility

A memory is worth recalling when conditioning on it makes the realized future more likely under the world model.

2Future-aware supervision

During training the observed future re-ranks the candidate memories, and the retriever learns to match that ranking without ever seeing the future at test time.

3Adaptive cue fusion

A learned gate weights time, pose, vision, audio, and any other cue per query, leaning on whichever is most trustworthy.

Experiments

Results

Three complementary settings, each with a different source of long-horizon uncertainty: LoopNav, a static outdoor Minecraft environment with time, pose, and visual cues; SoundSpaces, realistic indoor scenes with spatialized audio; and AI2-THOR, interactive household scenes whose state changes with the agent's actions.

Every trajectory has an exploration phase, in which the agent observes and, where applicable, interacts with the environment, and a return phase, in which it revisits explored locations. The world model generates the return-phase video using the exploration-phase observations as its memory pool, and we measure frame-wise discrepancy against ground truth.

Example evaluation trajectories

These are ground-truth episodes from the evaluation sets, not model rollouts. Each shows the exploration phase, which fills the episodic memory, followed by the return phase that the world model is asked to predict. Generated rollouts are shown further below, per dataset.

LoopNav (ground truth). Exploration, then return.
SoundSpaces (ground truth). Exploration, then return. 🔊 Has sound. Click the speaker icon in the controls to hear the spatialized audio the agent uses as a retrieval cue.
AI2-THOR (ground truth). Exploration with object interactions, then return.

LoopNav

FAR outperforms hand-designed recall

A static outdoor Minecraft environment with procedurally generated navigation, where retrieval can use time, pose, and visual cues. Each block of the reel shows the ground-truth frame, then for FAR, Temporal, WorldMem, and LongLive-RAG the prediction and the four recalled memories with their time offsets. Red shading marks where a prediction differs from the ground truth. It is an error overlay drawn for visualization, not something the model generated. FAR reaches 20 to 40 seconds back for the frames that matter, while the baselines mostly recall the last few seconds.

DreamSim at loop closure ↓ perceptual distance to the ground truth when the agent returns to its starting point; lower is better

0.200
Temporal
0.133
LongLive-RAG
0.130
WorldMem
0.082
FAR (ours) · 37% lower than the best baseline
Show full LoopNav rollouts (4 uncut episodes)

Full return-phase rollouts. Press play on a video to watch it; use the controls to scrub or go fullscreen.

LoopNav, episode 12
LoopNav, episode 1656
LoopNav, episode 2259
LoopNav, episode 3106

SoundSpaces

FAR adaptively learns which cues to trust

Realistic indoor navigation with spatialized audio. Each block shows the ground-truth frame, then for each method the prediction and the recalled memories with their time offsets. Red shading marks where a prediction differs from the ground truth. It is an error overlay drawn for visualization, not something the model generated. Where pose and appearance are ambiguous along a corridor, FAR learns to rely on audio and recalls a memory with a clear line of sight to the goal region.

DreamSim near loop closure ↓ perceptual distance to the ground truth over the final quarter of the return phase; lower is better

0.452
Temporal
0.423
WorldMem
0.150
FAR (ours) · 65% lower than the best baseline
Show full SoundSpaces rollouts (4 uncut episodes)

Full return-phase rollouts. Press play on a video to watch it; use the controls to scrub or go fullscreen.

SoundSpaces, episode 84
SoundSpaces, episode 99
SoundSpaces, episode 212
SoundSpaces, episode 1115

AI2-THOR

FAR recalls the right state as the world changes

Interactive household scenes where the agent moves objects between surfaces and containers during exploration, so memory holds the same place in conflicting world states. Each block shows the ground-truth frame, then for each method the prediction and the recalled memories with their time offsets; a badge on each prediction says whether the rendered state matches the world the agent left behind. FAR recalls the frames from the interaction itself rather than a pose-matched but stale view.

State-change prediction accuracy ↑ fraction of revisited scenes where the world model renders the changed state rather than the stale pre-interaction one; higher is better

Temporal
WorldMem
FAR (ours)
Surface changes
5.7%
42.1%
80.2%
Container reveals
0.6%
12.5%
75.7%

FAR is 1.9× the best baseline on surface changes and 6× on container reveals, where the contents only become visible after opening.

Show full AI2-THOR rollouts (3 uncut episodes)

Full return-phase rollouts. Press play on a video to watch it; use the controls to scrub or go fullscreen.

AI2-THOR, episode 309
AI2-THOR, episode 350
AI2-THOR, episode 430

Off-scene dynamics prediction

A controlled two-agent corridor: the observer sees a second, independently patrolling agent only intermittently through a doorway, and must later predict where that agent is when looking down the corridor. Each recalled context is a pair of observations about 1.5 seconds apart so it can carry motion.

Predicting the patrolling agent’s future location ↑ trajectory-level accuracy; higher is better

41.9%
Temporal
32.4%
WorldMem
67.1%
FAR, time & pose
94.6%
FAR, + agent cue
Off-scene dynamics, episode 47. Return-phase rollout in the two-agent corridor. FAR recalls a motion-informative episode containing the second agent and predicts where it will be; WorldMem retrieves a spatially relevant but dynamically uninformative memory.

Discussion and limitations

FAR learns both which episodic memories are useful for prediction and which retrieval cues to trust. It uses the observed future during training to assign predictive credit while remaining future-blind at inference, and its adaptive cue fusion lets it exploit time, pose, vision, audio, or any other signal attached to a memory. Across navigation, multimodal, and changing-state environments, this consistently improves long-horizon prediction and supports predictive utility as a principled basis for episodic memory access in persistent world models.

Our work focuses on the recall problem and assumes an external episodic memory. How memories are written, compressed, or forgotten, and how interactions among multiple memories are credited beyond our candidate-level approximation, remain open. FAR relies on informative retrieval cues and adds training-time computation for future-aware utility evaluation, though inference is unchanged. Our experiments are in controlled simulation; extending predictive episodic recall to richer real-world settings, together with learned memory formation and persistent-state representations, is a natural next step.

BibTeX

@article{kim2026learning,
  title   = {Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models},
  author  = {Kim, Beomsu and Lai, Chieh-Hsin and Nguyen, Bac and Bar, Amir and Ye, Jong Chul and Mitsufuji, Yuki},
  journal = {arXiv preprint arXiv:2609.34677},
  year    = {2026}
}