Source-linked AI summary
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
Sarthak Singh
TL;DR
DreamBench-SWE addresses how multi-session software agents retain and use non-inferable earlier-session evidence. It introduces an executable memory-hygiene benchmark and reports an original v2 fold plus a preregistered v2.1 successor audit, supporting benchmark discrimination and one bounded hosted-memory profile without establishing mechanism, superiority, equivalence, or broad product generality.
Problem
DreamBench-SWE studies memory hygiene as a benchmarkable failure mode when later software tasks depend on hidden, non-inferable evidence from earlier sessions.
Method
The paper uses hand-curated multi-session memory traps scored by executable oracles, alongside a training-free reference probe and controlled memory-system comparisons.
Results
The v2 primary comparison was null, while the v2.1 audit separated available memory conditions from no memory and characterized one bounded hosted literal-storage configuration without establishing mechanism or superiority.
Takeaways & Limitations
DreamBench-SWE is supported as a discriminating executable profile benchmark, with the successor audit providing one bounded external-system profile.
Takeaways & Limitations
HarmfulMemoryRate was 0.000 for every condition in the confirmatory fold, so the reported hygiene panel does not support harmful-retrieval reduction claims.
Abstract
from arXiv · showhide
DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. The successor run completed 360/360 work units and 720/720 S3 cells across four conditions. In the original fold, the primary DF-hybrid--B5 contrast was null (95/180 versus 89/180; clustered p=.518, Holm p=1), not evidence of equivalence, and C9/C10 retained B0-headroom limitations. In the successor, no external memory achieved 21/180 passes (rate 0.1167), deterministic verbatim event memory 82/180 (rate 0.4556), the typed-plus-raw reference probe 83/180 (rate 0.4611), and one pinned hosted Mem0 literal-storage configuration 97/180 (rate 0.5389). The registered six-slot Family A retained unavailable slots at p=1; all three available comparisons against no memory rejected after Holm correction. Both preregistered mechanism contrasts were unavailable after pre-evaluation conformance rejection. The secondary literal-storage-versus-verbatim comparison was nonconfirmatory and sensitivity-dependent, while the comparison with the reference probe did not reject. The audit therefore supports DreamBench-SWE as a discriminating executable profile benchmark and characterizes one exact hosted-memory configuration, but it does not establish an external-system mechanism, superiority among memory-bearing conditions, equivalence, or broad product generality. The original v2.0.5 findings and artifacts remain unchanged.
1 Introduction
DreamBench-SWE evaluates memory hygiene in multi-session software tasks where hidden earlier-session evidence must support later executable code outcomes. The paper reports a null original primary comparison and a successor audit that separates available memory-bearing controls from no memory while preserving narrow claim boundaries.
- Benchmark and motivation: Memory quality is not memory volume because stale retrievals, overgeneralized feedback, and uncertain failure diagnoses can harm later agent behavior.
- Benchmark and motivation: DreamBench-SWE uses hidden, non-inferable earlier-session evidence and executable oracles to benchmark memory hygiene in multi-session software tasks.Traps include scoped feedback, superseded architecture facts, generated-file boundaries, flaky failures, and exact-token contracts.
- Research questions: The paper’s central questions cover benchmark exposure of memory failures, hybrid-versus-verbatim ordering, structured-memory comparisons, and executable hygiene diagnostics.
- Original fold: 95/180 versus 89/180 was the original v2 primary reference-probe hybrid--B5 result, with clustered p=.518, Holm p=1.0, and reject=false.This is a failure to reject, not evidence of equivalence.
- Claim boundaries: The benchmark remains the headline contribution, whereas the v2 fold narrows anti-hoarding claims because C9 and C10 fail B0-headroom criteria.
- Successor audit: The v2.1 audit separated concurrent positive controls and one admitted hosted literal-storage configuration from no memory, while mechanism contrasts were unavailable after conformance rejection.The supported claim is benchmark discrimination plus one bounded external-system profile, not mechanism, ranking, superiority, equivalence, or product generality.
2 Related Work
DreamBench-SWE is positioned among software-agent, memory, and research-agent benchmarks without claiming generic priority. Its narrower contribution is executable coding-agent memory-hygiene evaluation with hidden code contracts and a bounded external-memory comparison.
- Benchmark positioning: Prior SWE benchmarks evaluate issue resolution, agent-computer interfaces, or harness optimization rather than longitudinal memory maintenance.
- Benchmark positioning: Conversational and agentic-memory benchmarks cover recall, temporal reasoning, updates, selective forgetting, dialogue, and action across varied domains.
- Benchmark positioning: DreamBench-SWE narrows its contribution to controlled repository continuation where hidden, non-inferable evidence must produce code edits passing executable repository oracles.The design includes CSPRNG token injection, scored-artifact isolation, and a taxonomy of SWE-specific pathologies.
- Memory architectures: The reference probe uses episodic, semantic, and procedural vocabulary as schema discipline rather than claiming cognitive novelty or ownership of existing memory frameworks.
- SWE-agent memory: The paper distinguishes its hidden exact-token, stale-fact, contradictory-feedback, and executable-oracle traps from subtask memory and persistent task-graph infrastructure.
- External systems and claim boundaries: Mem0 is the only external memory-system family represented as a scored condition; other related systems are positioning references rather than practical baselines.The successor admits only the pinned hosted literal-storage configuration.
3 Problem Formulation
The paper formalizes multi-session software-agent memory as append-only trajectories plus derived, scoped memory items whose admission and usefulness determine hygiene. Its central question concerns offline maintenance and raw-evidence retention under fixed agent and task conditions.
- Objects and sequences: A session is a bounded agent run, and a sequence is an ordered list of sessions over one repository or a controlled related-repository family.
- Trajectories and memory: Each trajectory records task metadata, events, outcomes, costs, and latency, while raw episodes preserve session evidence for later memory operations.
- Trajectories and memory: Derived memory items carry content, type, provenance, scopes, lifecycle status, validity, confidence, utility, risk, staleness, and relationship metadata.
- Memory hygiene: Read admissibility separates candidate generation from admission, rejecting memories whose status, validity, scope, supersession, risk, staleness, provenance, or type policy is unsuitable.
- Memory hygiene: Memory hygiene measures retrieval of useful, grounded, current, scoped, low-risk memories while suppressing harmful or stale items, rather than maximizing generic recall.
- Metrics and research question: HarmfulMemoryRate was 0.000 for every confirmatory condition, so the fold does not support harmful-retrieval reduction claims.
- Metrics and research question: RQ0 asks what offline maintenance, raw-evidence retention, and their combination change under fixed model, token, tool, and task-stream conditions.The descriptive ladder includes A0 episodic-only at 20/180 = 0.111 versus typed-only at 80/180 = 0.444.
4 Method: Reference Probe
The reference probe wraps a fixed coding agent with external memory, append-only raw evidence, offline maintenance, and gated retrieval. Its evaluated ladder separates typed, raw, and hybrid memory channels while preserving provenance and limiting causal claims about individual operators.
- System architecture: The reference probe adds an external memory store and offline maintenance loop around a fixed wake-phase software-engineering agent without training a new base model.
- System architecture: Each session reads condition-specific memories, runs the agent, logs the trajectory, and applies a sleep schedule for offline maintenance.
- System architecture: Raw episodes remain append-only evidence, while sleep operators may append derived memories, audit records, or lifecycle transitions but cannot mutate raw episodes.
- Evaluated variants: The frozen repairs add a verbatim raw-evidence capsule and exclude contradiction records from implementation reads through the allowed-types gate.
- Evaluated variants: Typed-only uses derived typed memories, raw-only exposes the raw-evidence capsule, and hybrid combines both channels.Raw-only is near-equivalent to B5 on the dominant recall-verbatim portion of the benchmark.
- Memory schema: Typed memory requires content, type, provenance, and write reason, with scope and lifecycle fields supporting repair and retrieval control.
- Retrieval gate: The retrieval gate applies hard admissibility filters before ranking candidates with semantic, scope, type, confidence, utility, verification, provenance, staleness, risk, age, and overscope terms.Weights, thresholds, candidate-pool size, and token budget are fixed before held-out evaluation.
- Evaluation scope: HarmfulMemoryRate was 0.000 for every condition, and several component ablations remained descriptive or designed-not-run rather than statistically isolated.
5 Experimental Design
The experiment evaluates DreamBench-SWE as a controlled multi-session benchmark using executable memory traps, predefined sequence types, and frozen condition and validity contracts. The confirmatory design compares memory policies under matched wake-phase budgets while retaining explicit limitations on baseline interpretation and reproducibility.
- Benchmark construction: 60 admitted three-session traps form the scaled v2 confirmatory fold, with S1 and S2 setup or reinforcement followed by an S3 memory-use test.Each S3 trap tests whether prior evidence is used without access to hidden oracles or reference solutions.
- Benchmark construction: The historical v1 set is unevenly distributed across five sequence families, so per-family results are descriptive and the benchmark is framed as a curated diagnostic stress test rather than broad production coverage.The v1 distribution includes 7 reviewer-preference, 5 convention-learning, 4 stale-architecture, 3 generated-files, and 3 flaky-test sequences.
- Benchmark construction: S3 traps require non-inferable evidence such as conventions, generated-file boundaries, stale architecture facts, reviewer preferences, or flaky-test lessons.Continuation permits a non-empty patch to advance even when its oracle fails, modeling ongoing projects with imperfect prior edits.
- Conditions: B5 is a deterministic verbatim event-memory substitute, whereas the reference probe uses typed and raw memory variants with lifecycle operations; neither comparison claims fidelity to a third-party system.B5 writes one deterministic memory per trajectory, while the reference probe writes multiple derived items and may mark them stale, superseded, or blocked.
- Controls: The successor audit froze its conditions, conformance gates, repair budgets, comparison families, and unavailable-slot rule before outcome inspection, with B0, B5, DF-hybrid, and hosted B5-MEM0-LIT clearing conformance.Native B5-MEM0 failed the normalized-context identity gate, while two local variants exhausted their repair budgets.
- Controls: The preregistered B0 warmup threshold was 0.80, but the canonical fold reports 287/360 = 0.797, so not every B0 S3 failure is interpreted as memory-only.The fold remains complete and its executable S3 outcomes remain reportable.
- Controls: Primary outcomes are Pass@1 and final_passed, supplemented by eight deterministic offline hygiene metrics with directionality defined separately for error-rate and accuracy measures.HarmfulMemoryRate, RepeatedErrorRate, RegressionAfterUpdate, ContradictionRepairAccuracy, HumanFeedbackUseAccuracy, UsefulMemoryPrecision, ScopeAccuracy, and TransferScore form the hygiene panel.
- Controls: Exact model behavior may not be independently reproducible because the pinned gpt-5.5 dependency is a hosted model that may be a mutable alias.The harness also pins the Codex CLI, node:22-slim base, image digest, and run commands.
6 Results
The v2 fold found a null primary comparison between the hybrid reference probe and B5, while the successor audit established discrimination from no memory but not superiority or mechanism claims.
- Probe ablations: 80/180 = 0.444, 84/180 = 0.467, and 95/180 = 0.528: typed-only, raw-only, and hybrid probe rates were descriptive rather than independently causal results.The planned marginal tests around these variants did not reject.
- Primary comparisons: 95/180 versus 89/180: the hybrid reference probe did not significantly separate from B5 in the primary clustered comparison.The clustered p-value was 0.518 and Holm-adjusted p=1.0; this was a failure to reject, not evidence of equivalence.
- Validity limitations: C9 and C10 were invalid anti-hoarding or abstention strata because B0 passed every cell, leaving no measured memory-dependent signal.C9 had 12/12 B0 passes and C10 had 6/6; these failures were not used to repair the null P1 outcome.
- Hygiene diagnostics: 1434/3024 = 0.474, 378/2268 = 0.167, 1512/2646 = 0.571, and 756/3780 = 0.200: executable hygiene metrics diagnosed benchmark-caught failure modes.These aggregate metrics were treated as diagnostics, not as a separate system-win family.
- Successor audit: B5 versus B0 and DF-hybrid versus B0 both rejected after Holm correction, while unavailable slots remained at p=1 in registered Family A.The successor audit therefore demonstrated benchmark discrimination against no external memory for the available comparisons.
- Successor audit: Neither registered mechanism contrast was evaluable because native B5-MEM0 and both Supermemory conditions failed frozen pre-evaluation conformance requirements.Unavailable conditions provide neither null results nor evidence of inferiority.
- Successor audit: 97/180 versus 82/180: B5-MEM0-LIT was numerically higher than B5, but the secondary comparison was nonconfirmatory and sensitivity-dependent.The comparison with DF-hybrid also did not reject, and the audit does not claim superiority among memory-bearing conditions.
- Interpretation: The successor supports DreamBench-SWE as a discriminating executable profile benchmark and characterizes one exact hosted literal-storage configuration.It does not establish an external-system mechanism, tuned-Mem0 or Supermemory performance, equivalence, or broad product generality.
7 Analysis
The v2 fold supports DreamBench-SWE as a validity-gated measurement protocol, but its primary system comparison was null and its external-system evidence remains bounded.
- Primary comparison: 95/180 versus 89/180: DF-hybrid numerically exceeded B5, but the clustered test did not reject (permutation p=.518, Holm p=1).This is a failure to reject, not equivalence.
- Family-level reading: All six preregistered clustered v2 comparisons were non-rejections after Holm correction, including typed-only, raw-only, hybrid, and the v1 structured-memory comparison.The system story is diagnostic and mechanistic, not confirmatory.
- Benchmark validity: 21/180: B0 was near floor overall, while strong memory baselines separated on valid C2, C3, C5, C6, and C7 strata.The fold produced complete condition-level result files and passed execution checks.
- Benchmark validity: C9 and C10 were not cleanly validated because B0 passed all cells in those strata, eliminating measured no-memory headroom.The remaining anti-hoarding constructs contributed 24 traps, above the 20-trap floor.
- Hygiene diagnostics: 0.264 versus 0.556: DF-hybrid’s UsefulMemoryPrecision was lower than B5’s, despite improved RepeatedErrorRate relative to B5.These hygiene metrics are diagnostics rather than standalone evidence of a system win.
- External-system audit: 21/180 and 20/180: the two pinned hosted Mem0 configurations remained near the no-memory floor in v2 and do not rank tuned Mem0 or Mem0g.The successor retained only one admitted hosted literal-storage configuration.
- External-system audit: The successor audit separated concurrent controls and one admitted hosted configuration from B0 after Holm correction, without establishing superiority, mechanism, equivalence, or product generality.It adds benchmark discrimination without revising the original null system comparison.
8 Limitations
The paper’s conclusions are bounded by implementation, benchmark, validity, transfer, infrastructure, and external-system scope limitations.
- Scope: The reference probe is not a new foundation model and does not establish that larger context windows are unnecessary or offline consolidation is always beneficial.The v2 fold shows that simple baselines can remain competitive with a richer maintenance pipeline.
- Inference: 95/180 versus 89/180 is a null, not an equivalence result; superiority or equivalence would require a new preregistration with suitable planning.The numerical direction should not be read as a hidden win.
- Benchmark validity: 287/360 = 0.797: B0 narrowly missed the 0.80 warmup gate, while C9 and C10 lacked no-memory headroom.These caveats constrain interpretation of memory failures and anti-hoarding or abstention claims.
- Memory pipeline: 0.264 versus 0.556: the typed retrieval pipeline failed the preregistered UsefulMemoryPrecision falsifier against B5.This limitation is independent of the null primary comparison.
- Transfer: Only one wake-model family was evaluated, and no full second-backbone transfer fold was reported.Broad model-generality language is therefore unjustified.
- External validity: Curated fixture repositories and controlled memory events may underrepresent production repositories, ambiguous intent, external services, and open-ended issue resolution.A future SWE-bench extension would answer a different transfer question.
- Evaluation setting: The evaluation is filesystem-isolated but not network-isolated because hosted model calls require network access.The guarantee rests on container boundaries and scanner/canary checks rather than an air-gapped environment.
- Metrics: Hygiene labels are deterministic diagnostics, but several are task-coupled or collinear and do not independently prove that the pipeline improves memory hygiene.Executable S3 pass/fail remains the primary outcome.
9 Conclusion
DreamBench-SWE reframes software-agent memory as a benchmarkable maintenance problem. The completed v2 fold yielded a bounded negative system result, while the additive v2.1 audit demonstrated benchmark discrimination and profiled one hosted configuration without establishing broader system claims.
- 9 Conclusion: DreamBench-SWE combines controlled repositories, hidden executable oracles, frozen sequences, injected non-inferable facts, validity gates, and trap-clustered inference.The evaluated maintenance pipeline serves as a reference probe combining raw retention with typed, provenance-linked maintenance.
- 9 Conclusion: 95/180 = 0.528 versus 89/180 = 0.494: the v2 primary comparison did not reject, and P2–P6 also did not reject.The correct interpretation is neither superiority nor equivalence.
- 9 Conclusion: C9 and C10 lacked B0 headroom, so they do not support spurious-lesson-rejection or abstention claims in this fold.These caveats narrow rather than invalidate the executed benchmark artifact.
- 9 Conclusion: 21/180, 82/180, 83/180, and 97/180: the successor audit separated all three available memory-bearing conditions from B0 after Holm correction.The registered mechanism contrasts were unavailable after pre-evaluation conformance rejection.
- 9 Conclusion: The paper supports a reproducible benchmark and one bounded external-system profile, not mechanism, broad model generality, tuned-Mem0 ranking, production reliability, or favorable cost claims.The evidence supports a transparent benchmark release with canonical artifacts for future systems.
A Recall-Verbatim Trap Audit
The recall-verbatim audit classifies most v1.0 traps as requiring prior event text, while distinguishing recall from copying and correcting denominator interpretation.
- Classification: 18 of 22 v1.0 S3 traps are classified as recall-verbatim under the evidence-location taxonomy.The classification follows the Q-BENCH semantic audit and preregistration disclosure.
- Classification: Four traps are synthesis/apply cases because their needed token is only partial without the generator workflow.All 22 traps still require implementation; none is a pure copy-and-stop task.
- Denominator: 44/65, not 44/66: the validity-gated typed-only denominator excludes one codex_exec_failed cell shown in the taxonomy display.
- Interpretation: Token-overlap fields are only crude code-diff signals; semantic classification asks whether prior injected event text contains the later hidden-oracle contract.B5 success is not automatic on every recall-verbatim row because implementation work remains necessary.
B Artifact and Reproducibility Details
The paper freezes reproducibility artifacts, protocols, condition definitions, and analysis records while distinguishing immutable v2.0.5 evidence from the successor audit.
- Frozen records: The canonical analyzer bundle and frozen sequence records provide the numeric source of truth for the v2 fold.The bundle includes clustered results, hygiene-oracle outputs, cost-frontier files, admission-funnel records, tables, and figures.
- Recorded inputs: The fold records fixed models, seeds, sequence files, and supplemental hosted-Mem0 rows, with B5-MEM0 and B5-MEM0-LIT excluded from P1–P6.The primary wake model is gpt-5.5 and complete conditions contribute 180 S3 cells each.
- Release boundaries: The public v2.0.5 artifact is immutable historical evidence, while the successor release is additive and does not alter it.The successor has its own canonical record and assets rather than revising the historical tag.
- Artifact contents: The package ships folded analyzer outputs and reproducibility scripts but excludes hidden scoring assets, credentials, logs, and full hosted-model result directories.It can verify reported paper numbers without live hosted-model reruns.
- Scope boundaries: External validity is bounded by filesystem-scoped isolation, unavailable transfer plans, and the absence of network isolation.The wake agent image also lacks Python and pytest, limiting direct repository execution.
- Protocol freezing: The v2 protocol was frozen before trap authoring and performance runs, with later changes requiring an amendment log.The fold fixes admission, analysis, falsifiers, and artifact standards before outcome inspection.
C.6 Primary Analysis
The primary analysis treats trap clusters, rather than repeated seed cells, as the inferential unit and precommits exact clustered testing, correction, validity gates, and claim constraints.
- Outcome and unit: The primary outcome is validity-gated S3 executable pass/fail for each condition, seed, and trap cell.Warmup, hygiene, cost, construct, family, and authoring measures are secondary or descriptive.
- Cluster statistic: For each trap, dt counts paired seeds won by A minus paired seeds won by B, with ties contributing zero.Only paired valid seed cells enter the comparison.
- Exact test: The primary p-value is a two-sided exact trap-clustered sign/permutation test with exchangeable signs and no normal approximation.Zero-contribution traps remain descriptive but do not add sign-flip states.
- Multiplicity and claims: Holm-Bonferroni correction controls family-wise error across P1–P6, and DF-hybrid > B5 requires a positive Dobs with Holm-adjusted p < 0.05.A null P1 cannot be rescued by pooled, slice, hygiene, or transfer analyses.
- Interpretive limits: The design has reasonable power only for large, stable trap-level effects and does not guarantee detection of small, sparse, or template-concentrated effects.The successor audit is separately frozen and cannot revise or pool with P1–P6.
E.5 v1 Clustered Sensitivity Result
The v1 clustered sensitivity analysis retained one narrow positive structured-memory comparison while showing that seed concentration and trap-majority choices materially affect sensitivity conclusions.
- Clustered result: The reference-probe typed-only versus B3 reflection-only comparison survived the v1 exact trap-clustered sensitivity check after Holm correction.It used 22 contributing trap clusters and 65 paired seed cells, with raw clustered p = 0.00732421875 and Holm-adjusted p = 0.03662109375.
- Sensitivity limits: The same comparison did not reject in all conservative sensitivity variants: per-seed tests and complete three-seed trap-majority collapse failed after correction.The trap-majority collapse had Holm-adjusted p = 0.1953125.
- Power unit: The v2 power analysis treats 60 traps as the effective clustered design rather than 180 independent cells.The relevant number of nonzero trap contributions and their seed clustering determines sensitivity.
- Thresholds: A 39–21 equal-weight trap split is the smallest split below 0.05, whereas 38–22 is not below 0.05.These thresholds use the frozen 180-cell denominator only for descriptive cell margins.
- Seed clustering: The same trap-level rejection boundary can represent a 10 percentage point or 30 percentage point cell-level margin depending on within-trap seed clustering.One-seed discordance per trap yields 18/180 = 0.100, while three aligned seeds yield 54/180 = 0.300.
- Power limitation: The design therefore does not guarantee detection of small effects, sparse discordance, or effects concentrated in a few trap templates.Its reasonable power is limited to large, stable trap-level reference-probe advantages.
F Reproducibility Appendix
The reproducibility appendix documents the repository layout, container boundaries, launch and validation commands, frozen inputs, result provenance, and a diagnostic hosted-Mem0 failure mode.
- Repository layout: The source tree separates paper text, benchmark code, frozen records, packaged artifacts, and live rerun machinery.Key areas include paper/, src/, scripts/, experiments/env/, experiments/results/, and dist/.
- Container boundary: Wake and judge execution uses container-filesystem isolation rather than network isolation, with hidden oracles and reference solutions excluded from the wake worktree.Default Codex and GRID image names and container-network overrides are documented.
- Frozen inputs: The confirmatory fold uses seeds 1, 2, and 3, with sequence records binding tasks to frozen repository commits and recording execution metadata.Result records link sequence source, oracle, harness, image, model route, and result directory.
- Secret protocol: Secret injection uses fresh cryptographic values without a deterministic replay flag, making rerunning injection over the same skeleton a new benchmark instance.Post-injection checks reject collisions, undeclared placeholders, and leakage outside allowed assets.
- Rerun semantics: The launcher is replay-ready but cannot recreate the historical fold exactly; reported completeness instead comes from canonical audits and folded records.Running it now creates a new hosted-model experiment.
- Hosted-memory diagnostic: The historical hosted-Mem0 configuration reached 6/66 = 0.091 pooled S3 pass@1, with sampled failures preserving retrieval but changing exact token bytes.A Unicode non-breaking hyphen replaced an ASCII hyphen, causing hidden-oracle rejection.
- Successor profile: In the successor comparison, B0 passed 21/180, B5 82/180, DF-hybrid 83/180, and B5-MEM0-LIT 97/180.The available comparisons against B0 were significant after Holm correction, while several registered conditions were unavailable.
G.6 Metrics and Hygiene Terms
The glossary defines task-coupled hygiene measures and frames paired or clustered analyses as sensitivity terminology rather than standalone headline outcomes.
- Hygiene values are deterministic offline re-scores from result records, treated as diagnostic because several metrics are task-coupled, collinear, or judge-dependent.
- RepeatedError Rate measures repeated labeled errors in active patches relative to opportunities to avoid repetition.
- StaleMemory Activation Rate measures stale or superseded memories used among tasks where stale memory was available.
- Contradiction Repair Accuracy measures correct repairs requiring the expected old-item status, new-item status, and relation.
- HumanFeedback Use Accuracy measures useful in-scope feedback retrieved and reflected in the active patch.
- Scope Accuracy measures the fraction of retrieved memories matching the task, repository, file, or symbol scope, combining feedback use and harmful-memory behavior.
H Ethics, Broader Impact, and Benchmark Integrity
DreamBench-SWE uses hidden synthetic evidence to resist public-artifact contamination while disclosing dual-use, hosted-service, computational, environmental, and configuration-specific boundaries.
- Hidden S3 literals are generated with a CSPRNG and omitted from public prompts, examples, artifacts, and recorded generator metadata.
- The anti-contamination claim is narrower than a general anti-cheating guarantee and does not establish network isolation or protect private reviewer packages after disclosure.
- The benchmark measures software-agent memory hygiene on synthetic repositories and events, with artificial tokens rather than credentials or private user and customer data.
- Hosted reference-probe configurations require explicit compute accounting, with diagnostic successful-task costs of 0.03832 USD for reference-probe hybrid and 0.03953 for B5.
- Sleep-phase maintenance materially increases model-call volume, limiting practical claims; future deployment studies should report budgets, call counts, runtime, and provider-backed energy estimates where available.
- The hosted Mem0 result concerns two pinned configurations under the reported harness, not Mem0 as a system family or tuned deployments.