Source-linked AI summary
Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision
Hongnan Zhao, Shiyu Chen, Zhihao Chen
TL;DR
Production content-generation systems need to coordinate immediate tasks, persistent brand identity, historical evidence, and revision feedback while preserving provenance. Aegix Pulse separates these responsibilities across a three-stage architecture and tests their incremental effects on synthetic social-media tasks. Account Profile and context-preserving revision showed encouraging but statistically inconclusive improvements, while successful-history evidence added no LLM-rated benefit and human validation was inconsistent.
Problem
Production generation must integrate current task requirements, persistent brand identity, historical evidence, factual sources, and revision feedback without treating them as one undifferentiated prompt.
Method
Aegix Pulse separates Task Persona clarification, versioned Account Profile assembly, successful-history evidence, controlled generation, and context-preserving revision in a preregistered component experiment.
Results
Account Profile and context-preserving revision produced encouraging improvements, but neither reached the prespecified significance threshold after multiple-comparison correction; successful-history evidence added no LLM-rated benefit.
Takeaways & Limitations
The findings support giving each context layer a clear purpose and provenance while preserving both positive and negative results for future production studies.
Takeaways & Limitations
The study used synthetic Chinese social-media tasks, offline content-quality evaluation, one hosted LLM Judge, and a small human review with poor reviewer agreement.
Abstract
from arXiv · showhide
Production content-generation systems must integrate a user's immediate task, long-term brand identity, historical evidence, and revision feedback. We present Aegix Pulse, a production-oriented three-stage architecture that separates current-task clarification and Task Persona finalization, long-term Account Profile (Brand DNA) assembly, and controlled generation and revision while preserving provenance across content versions. We evaluate four preregistered claims using 96 synthetic social-media generation tasks. Four initial-generation conditions progressively introduced a Task Persona, Account Profile, and successful-history style evidence, while two revision conditions compared plain and context-preserving revision. The experiment produced 480 completed generation records and 1,440 blinded LLM-Judge evaluations, supplemented by human review. Adding the Account Profile increased mean brand-consistency scores by 0.1562 points on a five-point scale compared with Task Persona alone (Holm-adjusted p=.1224). Preserving task and brand context during revision increased mean task-preservation scores by 0.2917 points compared with plain revision (Holm-adjusted p=.2432). Neither improvement was statistically conclusive after multiple-comparison correction. Task Persona alone showed a small observed effect, while successful-history evidence provided no additional improvement in brand consistency under the current setting. Human validation did not consistently reproduce the LLM-Judge effect directions and showed low inter-reviewer agreement. These findings provide preliminary evidence for persistent brand context and context-preserving revision while identifying priorities for stronger evidence processing and evaluation.
1 Introduction
Production content generation must coordinate immediate task requirements, persistent brand identity, historical evidence, and revision feedback without losing provenance. Aegix Pulse addresses this by separating these inputs and evaluating their incremental contributions.
- Motivation: Production use requires coordinating current goals, persistent brand position, successful prior expression patterns, factual sources, and narrow revision feedback.Treating these inputs as one undifferentiated prompt complicates evidence prioritization, revision state preservation, and output traceability.
- Architecture: Aegix Pulse separates the Task Persona, versioned Account Profile or Brand DNA, and successful-history style evidence across a three-stage workflow.Historical evidence supplements rather than redefines the Account Profile, while local revisions remain distinct from task redefinition.
- Evaluation: The study adds context components incrementally, comparing each condition with the immediately preceding condition across four preregistered questions.The questions address Task Persona alignment, Account Profile brand consistency, successful-history evidence, and context-preserving revision.
- Architecture: The architecture keeps current task, long-term identity, historical evidence, controlled generation, validation, and revision history clearly separated and traceable.This design supports measuring how individual context layers contribute to output quality.
- Evaluation: Account Profile and context-preserving revision produced encouraging preliminary improvements, whereas successful-history style evidence added no LLM-Judge improvement under the current setting.The paper reports both observed improvements and limits of the statistical evidence without expanding claims after seeing results.
2 Related Work
Prior work establishes personalized generation, retrieval-augmented evidence, memory systems, iterative refinement, and LLM-based evaluation. Aegix Pulse combines these ideas with explicit separation of task intent, brand identity, historical style evidence, and revision context.
- Personalization and memory: Personalization research distinguishes adapting outputs to a specific user from role-playing a fictional or predefined character.PersonaLLM demonstrates prompted personality expression, while related work emphasizes evaluating whether traits are represented accurately.
- Personalization and memory: LaMP retrieves user-profile items for personalized generation, while Aegix Pulse separately models current task intent, versioned Brand DNA, and successful-history style evidence.This three-way distinction directly motivates the stepwise E1, E2, and E3 design.
- Personalization and memory: Retrieval-augmented generation and hierarchical memory systems separate model knowledge from externally retrieved information and support traceable multi-session interaction.Aegix Pulse assigns distinct responsibilities to relational records, semantic indexes, and eligibility-filtered historical style evidence.
- Revision and evaluation: The revision experiment compares plain revision with revision that preserves the original Task Persona, Account Profile, style evidence, and generation context.Both conditions use identical source content and feedback, while Recipe and Validator/Repair remain disabled.
- Revision and evaluation: LLM-based evaluation suits open-ended generation but may favor model-generated text, so Aegix Pulse fixes prompts, hides labels, randomizes outputs, repeats judgments, and adds human review.The LLM Judge was relatively consistent, yet its effect directions did not always match human-reviewed results.
- Revision and evaluation: Preregistration reduces experimental overfitting by fixing hypotheses, procedures, evaluation rules, and analysis plans before formal results are examined.The study also locked the formal dataset and retained failed attempts rather than replacing them.
3 The Aegix Pulse Architecture
Aegix Pulse is a traceable architecture that assigns distinct roles to task clarification, persistent brand identity, historical style evidence, controlled generation, and revision lineage. Its engineering capabilities are deliberately distinguished from experimentally demonstrated quality effects.
- 3.1 Design principles: Aegix Pulse keeps information with different lifetimes separate and records how each generated version was produced.The design separates current task, long-term Account Profile, successful historical evidence, authoritative records, retrieval indexes, and audit history.
- Stages 1–2: Stage 1 clarifies requirements through a session and finalizes a Task Persona covering audience, topic, pain points, value proposition, content goal, and audience stage.The finalized Task Persona is linked to the exact Account Profile version and source materials used.
- Stages 1–2: Stage 2 combines the active Brand DNA with successful historical content only when a versioned success rule and sufficient performance evidence make retrieval eligible.Retrieved items retain performance and provenance metadata, and extracted tone, pacing, rhetorical habits, and presentation styles remain separate from both Persona objects.
- 3.3–3.4 Stage 3: Stage 3 builds Generation Context from the finalized task, exact profile version, memories, style evidence, effective style, and source materials before controlled candidate generation.Primary and supporting Recipes are recorded separately, while two candidates differ only in one defined variable.
- 3.4 Stage 3: Validator and Repair check requirements and factual, structural, and formatting constraints, but Recipe, Validator, and Repair were disabled in the formal experiment.Accordingly, the experiment does not attribute observed content-quality effects to these components.
- 3.5 Revision: Local revisions receive new linked Generation IDs, whereas changes to audience, topic, value proposition, content goal, or audience stage trigger a new clarification and Persona-finalization cycle.Before revision, unchanged and modifiable information is recorded; task-definition changes return RELOCK_REQUIRED.
- 3.6 Evidence boundaries: Relational databases store authoritative entities and histories, while isolated vector indexes support semantic retrieval without exposing another user’s private data.Engineering evidence supports traceable clarification, constraint checking, revision history, account isolation, and recoverable execution, not automatically better content or business outcomes.
4 Methods
The preregistered methods isolate four incremental context comparisons using controlled synthetic social-media tasks and predefined analysis procedures. Initial-generation and revision conditions differ only in the specified context additions.
- Hypotheses: The preregistered hypotheses test Task Persona alignment, Account Profile brand consistency, successful-history evidence, and context-preserving revision.H1–H4 compare E1>E0, E2>E1, E3>E2, and R1>R0 respectively.
- Experimental conditions: The four initial-generation conditions use the same model, decoding settings, schema, task, and factual source material, differing only in added context.Table 2 defines the E0–E3 ablation conditions.
- Experimental conditions: The E3 terminology denotes successful-history style evidence extracted in advance from synthetic successful posts and kept separate from the Account Profile.The terminology change from the original identifier does not affect the experiment, outputs, or statistical results.
- Revision conditions: The revision comparison gives R0 only the E3 output and feedback, while R1 additionally receives the original Task Persona, Account Profile, style evidence, and generation context.Both produce one revised version, with Recipe and Validator/Repair disabled.
- Dataset: The formal dataset contains 96 synthetic social-media tasks, with 48 selected for a balanced revision experiment.It spans 12 Account Profiles, six content goals, four audience stages, and 12 business scenarios.
- Dataset: Each task includes checked task and brand context, constraints, source material status, synthetic successful posts, extracted style evidence, and a local revision request.The synthetic posts are labeled experimental_verified_success and remain separate from production memory and performance data.
- Reproducibility: The formal dataset was locked with a SHA-256 hash after pilot and development cases were excluded from results.The hash permits verification that the dataset was unchanged after the experiment began.
4.4 Models and frozen prompts
The study froze model, prompt, evaluation, randomization, and analysis choices, then combined blinded repeated LLM judging with independent human review and paired statistical tests.
- Models and frozen prompts: DeepSeek-V3 generated content, while Qwen3.5-122B-A10B performed blinded evaluation; retrieval performance was not evaluated.The production embedding model was BAAI/bge-m3, and the hosted generator lacked an immutable revision identifier.
- Models and frozen prompts: The experiment manifest versioned the generator, revision strategy, Judge prompt, schema, aggregation policy, decoding settings, and randomization seeds.Candidate order was randomized within each task and replication, and the Judge received no condition labels, prompts, model identity, repair counts, or logs.
- Outcomes: Primary outcomes measured task-intent alignment, brand consistency, and revision preservation.Task-intent alignment covered audience, problem, topic, value proposition, goal, and audience stage; brand consistency covered profile alignment, positioning, tone, wording rules, and expression boundaries.
- Outcomes: Revision preservation combined five dimensions: task definition, unchanged content, requested change, facts, and brand style or structure.Additional operational checks covered factual and constraint violations, formatting, generation success, latency, retries, token usage, and estimated cost, but these were excluded from the confirmatory 1–5 Judge tests.
- Automated and LLM evaluation: Rule-based checks verified schemas, formatting, prohibited expressions, numerical claims, revision structure, and version history, while blinded LLM judging assessed factual and semantic compliance.The Judge was not shown the experimental condition for each output.
- Automated and LLM evaluation: Each candidate-context pair received three independent Judge evaluations, with preregistered medians and task-level aggregation producing initial-generation and revision scores.Revision used the lowest dimension score as its overall preservation score; unstable dimensions were routed to human review.
- Human review and adjudication: Two reviewers evaluated a blinded, balanced subset of 24 tasks covering all profiles, goals, stages, and conditions.They reviewed 96 initial-generation and 48 revision records; major score disagreements and disputed violations received third-person review.
- Statistical analysis: Conditions were compared within task using paired permutation tests, paired Cohen’s dz, Wilcoxon checks, bootstrap confidence intervals, and correction for the jointly tested hypotheses.The analysis reported means, standard deviations, medians, quartiles, and 95% confidence intervals from 10,000 task-grouped bootstrap resamples.
5 Results
The preregistered experiment completed all planned records and evaluations, with descriptive gains for account context and context-preserving revision but limited statistical and human-validation support.
- Execution: 480 generation records and 1,440 blinded Judge evaluations were completed successfully, with all planned outputs and audit fields preserved.All 384 initial-generation and 96 revision runs completed; 864 candidates passed SEO-format and prohibited-expression checks.
- Preregistered comparisons: Adding the Account Profile produced the largest initial-generation improvement, but it was not significant after multiple-comparison correction.The Task Persona showed a small task-alignment improvement, while successful-history style evidence produced a small negative brand-consistency difference.
- Descriptive condition scores: Mean task-intent scores for E0–E3 were 4.1667, 4.2292, 4.2865, and 4.2344, respectively, while brand-consistency scores were 4.0885, 4.1771, 4.3333, and 4.2760.Baseline scores were already high, leaving limited room for improvement on the five-point scale.
- Descriptive condition scores: Mean revision-preservation scores were 4.4583 for R0 and 4.7500 for R1, while feedback execution was 4.5000 versus 4.7917 and fact preservation was 5.0000 for both.Figure 4 presents these condition means descriptively; the preregistered comparison is reported separately.
- Judge stability: 99.09% of task-intent, 97.27% of brand-consistency, and 97.92% of feedback-execution score triplets differed by no more than one point across repeated LLM evaluations.Quadratic-weighted agreement scores were 0.7093, 0.7768, and 0.8220, respectively.
- Human validation: Human-review effects differed from LLM directions for most hypotheses, and first-reviewer agreement was poor with quadratic-weighted kappa of 0 for task intent and brand consistency.Only H4 had the same direction in human and LLM evaluations; 26.25% of adjudicated human-reviewed candidates contained a factual or constraint-related issue.
- Operational results: The estimated total model cost was CNY 28.508254, with median Judge latency of 8,273.5 ms and recovery from 19 rate-limit events and nine incomplete responses.No formal experimental record was deleted or overwritten.
6 Discussion
The discussion interprets the modest empirical effects alongside the architecture’s traceability and evaluation contributions. Account Profile context and context-preserving revision were the clearest positive results, but neither was statistically conclusive after correction.
- None of the four hypotheses reached the prespecified significance threshold after multiple-comparison correction.
- The Account Profile showed a clearer positive result than Task Persona alone, which may add limited value when the original request is already detailed.The Account Profile supplies long-term tone, positioning, and wording rules that may be absent from the current request.
- Successful-history style evidence did not further improve brand consistency beyond the Task Persona and Account Profile in this experiment.Its value may depend on real performance data, retrieval accuracy, available history, and domain.
- Preserving original task and brand context during revision may help retain important information during small changes, although the improvement was not statistically conclusive.This supports separating local Stage 3 revision from a new Stage 1 cycle when the task definition changes.
- Initial-generation scores above 4.0 on a five-point scale left limited room for improvement, making small effects harder to detect.The capable current-model baseline was realistic but reduced available headroom.
- The architecture and evaluation method are the paper’s main contributions, with separated context, provenance, revision history, and reproducible component testing.Engineering tests establish workflow properties, whereas effectiveness experiments test whether specific components improve measured outcomes.
7 Limitations
The study’s evidence is bounded by synthetic, offline, narrowly scoped experiments and evaluation limitations. These constraints limit generalization and the strength of conclusions about production effectiveness.
- The experiment used synthetic Chinese social-media tasks for cross-border e-commerce sellers and did not establish applicability to real users, other platforms, languages, or changing brands.It also omitted exposure, engagement, conversion, revenue, retention, and satisfaction outcomes.
- The formal evaluation relied on one hosted LLM Judge whose possible systematic bias and provider-side model updates limit exact reproducibility.Human-reviewer agreement was poor, and resolving disagreements did not improve reliability of the original ratings.
- Scores clustered near the top of the five-point scale, leaving limited room to detect improvement and potentially hiding minor differences.The revision aggregate could also be substantially reduced by one weak dimension.
- The study tested prechecked context components rather than automatic extraction, retrieval, successful-content identification, or style extraction in the complete production workflow.Recipe and Validator/Repair were disabled in the formal conditions.
- The formal dataset cannot be reused as independent evidence after its results became known.Any improved system or evaluation protocol requires a separate dataset and new preregistration.
8 Future Work
Future work should test the most promising components on independent, more demanding data and strengthen evaluation design. It should also isolate structural planning and validation effects in separate experiments.
- A new preregistered study should retest Account Profile context and context-preserving revision on a dataset created before outputs are examined.It should include more difficult and conflicting task requirements and better-calibrated human reviewers.
- Future production research should test whether style evidence from real successful posts improves brand consistency and user acceptance over time.The analysis should distinguish simple historical-performance correlations from stronger evidence.
- Recipe-based structural planning should be evaluated separately for structural fit, requirement compliance, and revision stability.Validator effectiveness should likewise be tested independently rather than mixed with other components.
- Future evaluations should combine rule-based checks, multiple independent LLM Judges, calibrated human reviewers, and outcomes reflecting actual user experience.Disagreements between evaluation methods should be reported and analyzed rather than collapsed into one score.
9 Conclusion
Aegix Pulse separates and traces task, brand, historical-style, generation, and revision context while evaluating components individually. The results are encouraging but preliminary: stronger effects were not statistically conclusive, and conclusions varied by evaluator.
- Aegix Pulse keeps current task, long-term Brand DNA, historical style evidence, controlled generation, and revision history separate and traceable.
- Adding the Account Profile and preserving original context during revision produced encouraging improvements, but neither reached the prespecified significance threshold after correction.
- Task Persona alone produced only a small improvement, while successful-history style evidence added no LLM-rated brand-consistency benefit in this experiment.
- The findings support purpose-specific context layers, explicit source priority, provenance records, and evaluation processes that preserve positive and negative results.They do not show that adding more personalization context always improves results.
Ethics, Data Governance, and Reproducibility
The study used synthetic, isolated experimental data rather than ordinary production-user content, with safeguards for review and publication. Reproducibility materials were retained, while future real-user studies are scoped to additional authorization, data minimization, access controls, retention rules, and ethical review.
- The study used synthetic tasks and kept experimental records separate from production memory, excluding ordinary production-user content, private messages, and real or simulated platform-performance data.
- Human-review files contained no account credentials, and generated content was not automatically published.
- The researchers retained the frozen protocol, dataset hash, definitions, Judge prompt hash, analysis scripts, audit results, and adjudication records for internal reproducibility.
- Future studies with real users or performance data will require explicit authorization, data minimization, access isolation, retention rules, and appropriate ethical review.
Competing Interests
The authors have a commercial interest because they are affiliated with the developer of Aegix Pulse. They addressed potential bias by fixing the protocol and analysis plan before examining formal results and reporting unsupported hypotheses openly.
- The authors are affiliated with Aegix Insight, which develops Aegix Pulse, creating a stated commercial interest in the system.
- To reduce potential bias, the experimental protocol and analysis plan were fixed before formal results were examined, and unsupported hypotheses were reported openly.