Beyond the TimelineAugmenting Long-Video Memory with Grounded Entity Biographies

Two red mugs in a kitchen. The striped mug used for coffee returns to the counter; a different solid red mug goes into the dishwasher. Two separate biographies answer the question.
Two red mugs, two biographies. A memory of moments cannot tell which mug went into the dishwasher; a memory of entities can.

Abstract

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the “biography” of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

Chronicle and biography

History can be organized around events or around the subjects who took part. A chronicle follows events through time; a biography follows a subject through those events. Long-video memory needs both: it must recover what happened at a particular moment and connect what happened to the same person or object across hours or days.

The same five moments, arranged twice. Read as a chronicle, two accurate descriptions of “a red mug” cannot say whether they concern the same mug. Read as biographies, the coffee mug never reaches the dishwasher.

Method

One blue hand mixer, followed from Day 1 to Day 6. The memory is written in three steps and read in three more.

  1. Ground
  2. Observe
  3. Associate
  4. Retrieve
  5. Read
  6. Answer
Day 1, 14:01. Jake holds a blue hand mixer at the sink while Shure reaches in to help. Day 1 · 14:01

described from its crops, the scene frames and the dialogue

Linked episode Shure explains “And then you put it on like this, oh.”
blue hand mixerone persistent physical instance
Day 3: the mixer beating cream at a crowded table.Day 3✓ match
Beating cream
Day 4: the mixer held by Shure.Day 4✓ match
Held by Shure
Day 5: the mixer mixing in a bowl.Day 5✓ match
Mixing in a bowl
Day 6: the mixer resting on the counter.Day 6✓ match
On the counter

An observation joins only if it matches the entity's recent references and is never seen apart from them in a shared frame.

Question While I was cleaning the egg mixer, who joined in to help?
Retrieval controller
search · “who helped me clean the egg mixer?”
Biography excerpt blue hand mixer

[Day 1 14:01] Shure shows how to attach the beaters.

Other recorded appearances 13:50, 13:55 …

Answer model“Shure”✓

Results

Four benchmarks over week-long and day-long recordings, with multiple-choice and open-ended questions. GEB has the highest overall score on every split.

Main results table: accuracy of sixteen systems on EgoLifeQA, Ego-R1-Bench and MM-Lifelong Test@Week and Test@Day. GEB reaches 72.0 on EgoLifeQA, 71.3 on Ego-R1, 36.83 on Test@Week and 17.58 on Test@Day, the best in each column.
EgoLifeQA and Ego-R1-Bench accuracy (%), MM-Lifelong GPT-5-judged accuracy (%). Bold: best; underline: second best. * and † mark published results; other entries are our runs.
MultiHop-EgoQA table: temporal grounding (mIoP, mIoG, IoU@0.3, mIoU) and answering (similarity, judge score) for whole-clip models and memory frameworks. GEB leads the memory frameworks on every metric.
MultiHop-EgoQA: temporal grounding and answer quality. GEB leads the memory frameworks on every metric and exceeds every reported model in IoU@0.3 and mIoU.

Ablation

EgoLifeQA accuracy (%) with one decision removed; controller and answer model fixed.

The indexAcc.The readingAcc.
GEB (full)72.0
w/o association68.6w/o biography text (index only)68.0
identity keyed by name69.2w/o biography text for the answer model70.8
descriptions appended to captions68.2w/o identity notes68.2
w/o observation→timeline edges68.2w/o unsearched-observation line70.2
w/o caption text61.8
w/o visual frames70.4

Evidence reaches the answer model more often

Bar chart: EgoLifeQA evidence hit rate rises from 37.6% with MAGIC-Video to 58.9% with GEB, in every question family.
EgoLifeQA: questions whose annotated evidence reaches the answer model, 37.6% to 58.9%.
Bar chart: MultiHop-EgoQA complete evidence coverage by number of annotated intervals; GEB's relative gain over MAGIC-Video grows with the number of intervals.
MultiHop-EgoQA: questions with every evidence interval reached. The gain grows with the number of intervals.
MultiHop-EgoQA evidence coverage table: fraction of questions with every annotated interval reached, retrieved duration, judge score and mIoU for MAGIC-Video, a six-unit MAGIC-Video variant, WorldMM, GEB, and two GEB ablations.
Complete evidence coverage on MultiHop-EgoQA. GEB reaches every annotated interval far more often than MAGIC-Video, and as often as WorldMM while reading less of the clip.

Two caps, two biographies

Question: who wore a hat while walking in the park. Two blue caps appear in one frame on Day 3. Cap 28648's biography names Lucia; cap 28919's biography names Tasha. Answer C: Lucia, Tasha.
Two blue caps share one frame at 15:09, so they are two objects. Each cap's biography reaches the episode that names its wearer.

Citation