Source-linked AI summary

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li

arXiv:2608.02218v1cs.AI

TL;DR

Existing poster-generation paradigms do not jointly provide request-level artifact validity, native editability, explicit design control, and measured per-request cost. PosterMELD addresses this with a template-conditioned multi-agent pipeline and achieves 81.3% PRR—3.4 times P2P’s and 5.2 times PosterGen’s rates—across 621 papers.

  • Problem

    Existing approaches do not jointly provide request-level artifact validity, native editability, explicit design control, and measured per-request cost.

  • Method

    PosterMELD uses capacity-aware template slots, five skill-guided agents, deterministic gates, VLM review, and bounded repair to produce controlled editable posters.

  • Results

    81.3% PRR—3.4 times P2P’s and 5.2 times PosterGen’s rates—was achieved across 621 papers, with the highest conditional CHE among generated methods.

  • Takeaways & Limitations

    Print readiness, native editability, and controllable design diversity can be achieved jointly in practical paper-to-poster generation.

  • Takeaways & Limitations

    Automatic gates certify geometry and legibility rather than scientific correctness, so authors must perform final verification.

Abstract

from arXiv · show

Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly. PosterMELD is a template-conditioned multi-agent pipeline: capacity-aware slots guide writing before rendering, and deterministic gates plus vision-language model (VLM) review route failures to bounded repair. Each accepted request exports editable PowerPoint (PPTX) and Portable Network Graphics (PNG) artifacts; explicit design controls yield same-paper variants. Across 621 papers, Print-Ready Rate (PRR) counts requests passing geometric, readability, asset-integrity, and obvious-factual-error checks, with native editability reported separately. A frozen VLM assigns conditional Craftsmanship-Harmony-Expressiveness (CHE) scores to print-ready outputs. PosterMELD attains 81.3% PRR, 3.4 times P2P's rate and 5.2 times PosterGen's, and the highest conditional CHE among generated methods with multiple print-ready outputs. Native editability and explicit design controls are retained at a mean cost of USD 0.38 per request, 3.5% of Codex+Skill's. Code and resources are available at https://github.com/Shannon4Science/PosterMELD.

Introduction

Scientific poster creation is labor-intensive and must satisfy geometric, density, revision, editability, and design-diversity requirements. PosterMELD addresses persistent print-readiness and controllable-diversity obstacles with template-first, capacity-aware composition, editable outputs, deterministic and VLM review, and bounded repair.

  • Motivation: Scientific posters require reorganizing claims, evidence, and visuals onto one constrained canvas, while revisions make native editability and explicit design controls practical requirements.Construction must accommodate venue, orientation, and density constraints before printing.
  • Persistent obstacles: Prior systems can hide request-level failures by conditioning aesthetic and content scores on generated outputs, while raster formats prevent element-level revision.Writing content before final geometry can cause shrinking, truncation, reflow, unreadably small text, or missing assets.
  • Persistent obstacles: Controllable design diversity remains difficult because stochastic reruns may produce omitted evidence, tiny text, or broken layouts, while coding-agent and direct-image workflows trade cost or editability against variation.The introduction identifies no existing paradigm as directly providing the evaluated combination of useful variation, native editability, and reliable text.
  • PosterMELD approach: PosterMELD fixes structure before writing content, using template-region capacity to condition keypoint selection, writing, and visual allocation on known geometry.Five skill-guided agents compose the poster, and each rendered draft undergoes deterministic gates and VLM review before failed aspects receive bounded repair.
  • Evaluation scope: 621 papers across 14 publication sources and ten domains form the largest end-to-end task-specific benchmark described, reporting request-level Print-Ready Rate alongside aesthetics, content, fidelity, and cost.For each request passing its gates, the pipeline produces an editable PPTX file and its PNG render while recording provenance.

Method

PosterMELD combines capacity-aware template contracts with a structured five-agent pipeline that plans, writes, lays out, renders, reviews, and repairs editable posters. Deterministic validation and bounded repair govern acceptance while explicit controls support independently requested design variants.

  • Template Library: Author-designed posters are mined into reusable templates through semantic block extraction, spatial layout descriptors, and Ward hierarchical clustering.Layouts use standardized geometry, block statistics, and an up-weighted 12×12 occupancy-grid segment; cluster-center samples become representative templates.
  • Capacity-Aware Templates: 4–10 content slots per template, with a mean of 6.1, encode geometry, reading rank, semantic role, text budgets, visual footprints, and compatibility constraints.Templates are augmented with region-level slot contracts so content is planned against available capacity before rendering.
  • Multi-Agent Pipeline: Five agents advance a shared typed poster state: four compose the draft, while a fifth renders, reviews, and repairs it until acceptance or budget exhaustion.Paper parsing preserves reading order and source locations, while keypoints and reused assets retain grounding and provenance links.
  • Design Controls: Users can fix template, style, density, logos, asset generation, and seed, while automatic mode selects a compatible template from paper statistics.An incompatible explicit template choice is reported rather than silently substituted, and design variants are requested independently by changing controls.
  • Review and Acceptance: Deterministic gates check artifact validity, geometry, readability, asset integrity, and occupancy, while VLM review inspects rendered posters and targeted crops for perceptual failures.The Review Agent combines structured checks with visual inspection rather than relying on coordinates alone.
  • Bounded Repair: Bounded repair applies scoped rewrite, reflow, resize, or rerender actions to failed aspects while preserving previously valid blocks.The Pipeline Harness freezes configuration and seed, enforces budgets, routes failures, validates artifacts, and records execution outcomes.

Benchmark and Evaluation

Benchmark and Evaluation establishes a 621-paper, cross-domain benchmark covering diverse publication sources and evaluates completion, print-readiness, aesthetics, content quality, and cost at the request level. The protocol separates output validity from pipeline completion and conditions CHE on print-ready posters.

  • Benchmark: 621 papers form the largest released end-to-end evaluation set, combining newly curated and re-annotated paper–poster pairs.The benchmark is 5.1 times the size of P2PEval and 6.3 times the size of the paired Paper2Poster set; every method runs on all requests.
  • Benchmark: The benchmark expands coverage across 14 publication-source groups and ten domains, including biology, medicine, psychology, and the social sciences.It also includes figure-heavy and equation-heavy papers, extending beyond core artificial intelligence coverage.
  • Baselines: Five baselines span dedicated pipelines, direct image generation, and coding-agent stacks, with shared GPT-4o execution for agentic pipelines.GPT-Image-2 and Codex+Skill are closed-source and reported separately; Codex+Skill uses GPT-5.4 with low reasoning effort.
  • Evaluation Metrics: Print-Ready Rate counts requested outputs that both complete and pass geometric, readability, asset-integrity, and obvious-factual-error checks.The shared request-level denominator prevents execution failures from being hidden by conditional quality scores.
  • Evaluation Metrics: 1,398 renders receive independent human print-readiness labels, requiring at least three positive annotations, while CHE is evaluated only on accepted posters.The human-labeled subset contains 1,398 available renders from 201 papers, and 661 posters are accepted by at least three of four annotators.
  • Evaluation Metrics: CHE averages 1–5 scores for visual craftsmanship, stylistic harmony, and expressive distinctiveness over print-ready posters.Unusable outputs are excluded from CHE and accounted for separately by PRR; content quality additionally follows P2P’s Universal protocol and keypoint-conditioned BERTScore.

Experiments

Across 621 papers, PosterMELD achieves high print readiness and conditional visual quality while preserving native editability at low cost. Qualitative results show controllable diversity, although external-model dependence and judge limitations remain.

  • Performance: 81.3% PRR across 621 requests exceeds P2P's 24.2%, PosterGen's 15.8%, and Paper2Poster's 0.2%.Using unrounded rates, this is 3.4 times P2P's PRR and 5.2 times PosterGen's.
  • Performance: 3.247 conditional CHE is highest among generated methods with at least two print-ready outputs, while PosterMELD's craftsmanship score is 3.455.Its craftsmanship score is closest to the human reference.
  • Evaluation: PosterMELD leads open-source systems on Universal, while GPT-Image-2 scores 4.948 on Universal but has the lowest conditional CHE, 2.698, among qualifying methods.Universal and BERTScore produce a different ordering from PRR and CHE; BERTScore ranks P2P at 0.829 above the human reference at 0.794.
  • Efficiency and editability: $0.38 per request preserves editable output alongside 81.3% PRR, compared with Codex+Skill's 82.8% PRR at $10.78 and GPT-Image-2's 85.2% PRR without editability.PosterMELD's mean request cost is approximately 28 times lower than Codex+Skill's.
  • Controllable diversity: Same-paper variants differ in presentation, organization, emphasis, and style while preserving shared research grounding, with PosterMELD varying templates, palettes, and text–figure allocation.Figure 5 compares outputs across five benchmark papers and four methods at a common rendered size.
  • Limitations: Parsing errors can propagate from unusual layouts, automatic gates do not certify scientific correctness, and frozen-judge biases are quantified but not eliminated.Final verification remains with authors, while calibration against human consensus and comparison across four judge families assess judge risk.

Related Work

Scientific poster generation has evolved from layout and content selection toward editable planning, agent–checker pipelines, aesthetic optimization, and post-hoc editing. PosterMELD integrates these directions with template-first capacity planning, bounded repair, explicit design controls, and request-level evaluation.

  • Automated scientific poster generation: Scientific-poster systems progressed from panel layout and content selection to editable planning, agent–checker pipelines, aesthetic optimization, and post-hoc editing.Poster-specific models coexist with diverse generation approaches.
  • Agentic artifacts and quality control: Slide-generation systems summarize sources, plan multimodal slides, reconstruct editable designs, and manipulate presentation objects.Document layout analyzers recover structural regions from heterogeneous pages.
  • Agentic artifacts and quality control: Poster systems add VLM repair, deterministic figure insertion, and cross-artifact coding skills, while verbal reflection supports iterative refinement.Agentic visual generation also introduces cognitive search, reasoning, and code-mediated canvases.
  • Agentic artifacts and quality control: PosterMELD contributes their integration with template-first capacity planning, bounded repair, explicit design controls, and request-level evaluation.The supplied passage identifies this integration as the contribution.

Conclusion

PosterMELD uses a template-first multi-agent pipeline to generate editable, print-ready scientific posters with explicit design controls. Capacity-aware contracts, deterministic validation, VLM review, and bounded repair preserve artifact validity and editability while producing matched deliverables and provenance.

  • PosterMELD uses capacity-aware slot contracts to guide five skill-specialized agents within fixed poster geometry.The template-first design guides generation before rendering.
  • Deterministic gates, VLM review, and bounded repair enforce artifact validity while preserving editability.
  • Each accepted request produces a native PPTX, matched PNG render, and provenance record.
  • Independent control settings provide design variation across generated posters.
Loading 2608.02218v1…