nanobot/nanobot/templates/agent/dream_phase1.md
chengyongru 35f3084c03 feat(dream): per-line age annotations + dedup-aware prompt + max_iter=15
Three improvements to Dream's memory consolidation:

1. Per-line git-blame age annotations: MEMORY.md lines get `← Nd` suffixes
   (N>14) from dulwich annotate. SOUL.md/USER.md excluded as permanent.
   LLM uses content judgment, not just age, to decide what to prune.

2. Dedup-aware Phase 1 prompt: reframed as dual-task (extract facts +
   deduplicate existing files) with explicit redundancy patterns to scan for.
   Validated through 20 experiments (exp-002 prompt + max_iter=15 was best,
   averaging -1643 chars/5.4% compression per run).

3. Phase 1 analysis as commit body: dream git commits now include the full
   Phase 1 analysis for transparency via /dream-log.

4. max_iterations raised from 10 to 15: 30% improvement over 10 with no
   risk; 20 showed diminishing returns (exp-020: -701 vs exp-017: -1643).
2026-04-17 13:45:38 +08:00

2.4 KiB

You have TWO equally important tasks:

  1. Extract new facts from conversation history
  2. Deduplicate existing memory files — find and flag redundant, overlapping, or stale content even if NOT mentioned in history

Output one line per finding: [FILE] atomic fact (not already in memory) [FILE-REMOVE] reason for removal [SKILL] kebab-case-name: one-line description of the reusable pattern

Files: USER (identity, preferences), SOUL (bot behavior, tone), MEMORY (knowledge, project context)

Rules:

  • Atomic facts: "has a cat named Luna" not "discussed pet care"
  • Corrections: [USER] location is Tokyo, not Osaka
  • Capture confirmed approaches the user validated

Deduplication — scan ALL memory files for these redundancy patterns:

  • Same fact stated in multiple places (e.g., "communicates in Chinese" in both USER.md and multiple MEMORY.md entries)
  • Overlapping or nested sections covering the same topic
  • Information in MEMORY.md that is already captured in USER.md or SOUL.md (MEMORY.md should not duplicate permanent-file content)
  • Verbose entries that can be condensed without losing information For each duplicate found, output [FILE-REMOVE] for the less authoritative copy (prefer keeping facts in their canonical location)

Staleness — MEMORY.md lines may have a ← Nd suffix showing days since last modification:

  • SOUL.md and USER.md have no age annotations — they are permanent, only update with corrections
  • Age only indicates when content was last touched, not whether it should be removed
  • Use content judgment: user habits/preferences/personality traits are permanent regardless of age
  • Only prune content that is objectively outdated: passed events, resolved tracking, superseded approaches
  • Lines with ← Nd (N>14) deserve closer review but are NOT automatically removable
  • When removing: prefer deleting individual items over entire sections

Skill discovery — flag [SKILL] when ALL of these are true:

  • A specific, repeatable workflow appeared 2+ times in the conversation history
  • It involves clear steps (not vague preferences like "likes concise answers")
  • It is substantial enough to warrant its own instruction set (not trivial like "read a file")
  • Do not worry about duplicates — the next phase will check against existing skills

Do not add: current weather, transient status, temporary errors, conversational filler.

[SKIP] if nothing needs updating.