MULTIMODAL MEMORY BENCHMARK

CUE-Mem

现有长期记忆评测多数只关注文本或显式多模态信息,对用户消息中反复出现的隐式线索关注不足:
1、图片背景、边缘
2、语音消息的背景音

本研究构建 CUE-Mem Benchmark,旨在评估多模态智能体从隐式线索中检索并应用长期用户记忆的能力。

20users
2,674QA pairs
648sessions
5,798dialogue turns
496images
2,228audio clips

DATA CONSTRUCTION

Benchmark data
construction process

01Profilepreferences
02Entityreferences
03Eventannotations
04Artifacthistory

Two representative cases illustrate how profile-level preferences are transformed into structured events and multimodal history.
Select a case, then use the step controls to reveal the construction sequence.

LIVE PIPELINE

Loading case…

准备播放
正在读取构造案例…

QA EXAMPLES

Four evaluation settings

Each example presents the question and retrieved memory clues first. Expand the answer panel to reveal the gold answer and the preference being evaluated.

正在读取 QA 案例…

EXPERIMENTAL RESULTS · DEMO V3

Three research questions

正在读取实验结果…
NOTE

SCOPE OF THIS DEMO

This section presents the benchmark's data construction process followed by four representative QA settings. The answer panels stay collapsed so the memory evidence can be inspected before the gold answer.