Probe: A Hop-Depth-Controlled, Bitemporal Diagnostic Dataset for Multi-Hop Contradiction Propagation in Conversational Memory
Description
This dataset is a diagnostic probe for evaluating contradiction propagation in bitemporal conversational memory for large language model (LLM) agents. It isolates variables that existing conversational-memory benchmarks do not control directly: the number of reasoning steps ("hops") separating a contradiction from the fact it should invalidate, the update-versus-correction axis of a bitemporal contradiction, and the ability to reconstruct a past state after later, possibly contradicting, events have been recorded. probe.json (168 items) is the core set. It crosses five design axes: hop depth (1 to 4), whether the contradicted edge is directly asserted or derived, whether the contradiction is an update or a correction, turn density (adjacent or distant across sessions), and stated confidence (high or hedged), realized across three domains (employment, residence, training). It has been fully human-validated by an independent annotator (168/168 PASS). probe_pit.json (8 items, 24 scored queries) is a point-in-time extension targeting bitemporal reconstruction specifically: given a conversation containing a double update or a correction followed by a later update, can a system correctly answer what was true as of an earlier reference time, not just the current state. It covers two domains (employment, residence) and has not yet been independently human-validated. All conversational turns in both files are produced deterministically by a fixed Python string-template generator. No large language model was used to write, paraphrase, or otherwise generate any conversational content. Gold answers are likewise computed programmatically from the same template values used to write each conversation, not inferred or verified by an AI system. Field reference, probe.json (per item): probe_id; domain, hop_depth, contradicted, axis, density, confidence (design-axis values); sessions (ordered conversation turns); question; gold_pre / gold_post (correct answer before / after the final turn; gold_post is "unknown" when the contradiction removes the basis for an answer without supplying a replacement); gold_invalidated (relations no longer trusted); gold_axis. oracle_facts and oracle_retractions are an internal structured representation used for a separate automated consistency check and are not required to interpret or grade an item. Field reference, probe_pit.json (per item): pit_id; domain, scenario (double_update or correction_then_update), density; sessions; question; queries (a list of evaluation points, each with a label, a reference_time, and the gold answer as of that reference time; every item includes a before_any_fact query, for which gold is "unknown" by construction). oracle_facts and oracle_retractions serve the same internal role as in probe.json.
Files
Institutions
- Ho Chi Minh City University of Foreign Languages and Information TechnologyHo Chi Minh, Ho Chi Minh City