Source-linked AI summary
Seed1.8 Model Card: Towards Generalized Real-World Agency
Bytedance Seed
TL;DR
Many real-world applications require interaction, tool use, and multi-step execution beyond single-turn prediction. Seed1.8 integrates these capabilities with strong LLM and VLM performance, configurable efficiency controls, and evaluations spanning benchmarks and practical workflows.
Problem
Real-world applications require models to use tools, incorporate environment feedback, and execute multi-step tasks rather than only make single-turn predictions.
Method
Seed1.8 combines LLM and VLM capabilities with unified search, code execution, GUI interaction, configurable thinking modes, and optimized visual encoding.
Results
Evaluations cover foundational capabilities, multimodal understanding, video, and agentic workflows across public benchmarks and practical use cases.
Takeaways & Limitations
Seed1.8 is released to support further experimentation and development for interactive, real-world use cases.
Abstract
from arXiv · showhide
We present Seed1.8, a foundation model aimed at generalized real-world agency: going beyond single-turn prediction to multi-turn interaction, tool use, and multi-step execution. Seed1.8 keeps strong LLM and vision-language performance while supporting a unified agentic interface-search, code generation and execution, and GUI interaction. For deployment, it offers latency- and cost-aware inference, including configurable thinking modes and optimized visual encoding for images and video. We report evaluations on standard benchmarks and application-aligned workflows spanning foundational skills, multimodal understanding, and agentic behavior. Seed1.8 is released to support further research and development on interactive, real-world use cases.
1 Introduction
Seed1.8 targets generalized real-world agency by combining strong language and vision-language capabilities with unified interaction, tool use, and multi-step execution. Its design also addresses practical deployment constraints through configurable inference and efficient visual processing.
- Motivation: Seed1.8 extends strong LLM and VLM capabilities beyond single-turn prediction toward interactive, multi-step task execution.The motivating setting involves tool use, environment feedback, and iterative execution.
- Model Design: The model integrates perception, reasoning, and action within a single model rather than relying on task-specific agent pipelines.This integration is intended to support generalized real-world agency.
- Agentic Interaction: Seed1.8 provides search, code generation and execution, and GUI interaction through a unified agentic interface.Intermediate retrieval, execution, and environment results inform subsequent actions.
- Deployment: Configurable thinking modes and optimized visual encoding balance inference depth, latency, computational overhead, and token consumption.The efficiency measures target interactive deployment with multimodal and long-context inputs.
- Evaluation: Development and validation combine public benchmarks with internal evaluations spanning foundational capabilities, multimodal understanding, and agentic workflows.The evaluation scope is intended to reflect realistic usage patterns.
2 Evaluation
Seed1.8 is evaluated across foundational language, multimodal, video, and agentic capabilities, with strong results across many benchmarks and practical workflows. The reported evaluations also examine efficiency, tool use, and deployment-oriented tasks.
- Foundational Capabilities: Seed1.8 performs competitively with leading models in coding, mathematics, STEM, general reasoning, instruction following, and broad knowledge.It achieves the second-highest scores on BeyondAIME, AMO-Bench, and IMO-AnswerBench, and the second-best score on Inverse IFEval.
- Image Understanding: Seed1.8 improves over Seed1.5-VL across visual tasks and approaches current state-of-the-art performance, including wins on several challenging benchmarks.The comparison includes Claude-Sonnet-4.5, GPT-5.1 (High), Gemini 2.5 Pro, Gemini 3 Pro, and Seed1.5-VL.
- Multimodal Reasoning: 11.0 Pass@1 on ZeroBench surpasses Gemini 3 Pro’s 10.0, while Seed1.8 also leads VLMsAreBiased with 62.0 versus 50.6 and ranks first on MUIRBench.Across 7 of 9 other multimodal reasoning benchmarks, it achieves the second-highest scores.
- Specialized Vision Tasks: 90.7 Pass@1 on DA-2K and 25.8 Pass@1 on MMSIBench establish new state-of-the-art results in 2D and 3D spatial understanding.These scores exceed Gemini 3 Pro’s 82.1 and 25.4, respectively.
- Video Understanding: Seed1.8 achieves state-of-the-art performance in 5 of 6 motion-and-perception tasks, but its TOMATO score of 60.6 remains below human performance of 95.2.The video evaluation also reports leading results in video knowledge and reasoning.
- Agentic Capabilities: 93.2 on GAIA exceeds GPT-5-high’s 76.7, while Seed1.8 also reports strong results in dedicated search, coding, tool-use, and writing evaluations.It scores second-best on AInstein-SWE-Bench and Terminal Bench 2.0 and performs on par with leading models on several other agentic benchmarks.
- Application-Aligned Workflows: Seed1.8 reports practical-workflow strengths, including 62.0 in Finance and 55.2 in Law on XpertBench and best performance on the multimodal WorldTravel benchmark.These evaluations target expert workload automation and daily-life planning.
3 Use Cases of Seed1.8
Seed1.8 is demonstrated on practical, multi-step tasks spanning travel planning, expert workflows, scientific research, and scientific software engineering. These examples require information synthesis, visual interpretation, domain knowledge, reasoning, tool use, and code execution.
- Travel planning: Travel planning requires navigating fragmented information across platforms while balancing time, budget, preferences, and dynamic visual interfaces.The WorldTravel benchmark uses synthetic webpages to mirror these everyday planning demands.
- Travel planning: Seed1.8 generates a full-day Berlin itinerary that satisfies family, budget, and other user constraints by synthesizing information from diverse web sources.The scenario uses travel aggregators, booking portals, restaurant menus, reasoning, tool use, and visual interpretation of web interfaces.
- Expert tasks: Seed1.8 handles expert-level professional workflows that require deep domain knowledge and complex, multi-step procedures.These examples are drawn from the internal XpertBench, with full responses reported in an appendix.
- Scientific research: Seed1.8 solves a biology research problem series from visual inputs, including questions about optogenetics, caspase structure, fluorescence, and immunoblot results.The example is presented as an internal BIOBench task and covers interpretation of experimental results and biological images.
- Cross-use-case pattern: The examples collectively illustrate real-world assistance that combines scientific knowledge, multimodal understanding, agentic coding, and multi-step execution.The scientific software task is framed as part of an internal benchmark for scientific software engineering and agentic coding capabilities.
- Scientific software engineering: In scientific software engineering, Seed1.8 completes missing functionality in an EinsteinToolkit numerical-relativity repository through mathematical derivation, numerical-stability analysis, and tool-assisted code exploration.The task involves implementing BrillLindquist while preserving specified metric handling, derivatives, and time-symmetric initial data.
4 Safety
Seed1.8 is evaluated for safety on open-source benchmarks and internal risky-content categories. The reported examples show refusals, legal or safety warnings, and responsible redirection across several domains.
- Open-Source Benchmarks: Seed1.8 improves substantially on AIR-Bench while maintaining high performance on XSTest.These are the two open-source safety benchmarks reported in Figure 10.
- Internal Safety Evaluation: The internal safety benchmarks cover Civil Norms, Pornography, Illegal Acts, Copyright, Medical Safety, and Identity.The evaluation focuses on risky-content categories and summarizes the model as consistently rejecting unsafe inputs.
- Risk Response Examples: The examples also include refusal of vulgar sexual content and clarification that DeepSeek is not affiliated with Doubao.These responses address civil communication norms and identity confusion, respectively.
- Risk Response Examples: For explosive-related requests, the model identifies TNT as controlled, warns about legal and physical risks, and refuses private manufacture.The response emphasizes criminal-law restrictions and extreme explosion and toxicity hazards.
- Risk Response Examples: For piracy requests, the model refuses unauthorized film links and directs users to legal viewing channels.The example recommends legal platforms or cinemas instead of sharing copyrighted resources.
- Risk Response Examples: For medical questions, the model provides conditional information about nifedipine and advises strict medical supervision without private dosage changes.The example notes contraindications including severe hypotension and acute myocardial infarction.
5 Conclusion
Seed1.8 is presented as a foundation model for generalized real-world agency, combining strong language and vision-language capabilities with multimodal, tool-use, and multi-step execution support. Its development and evaluation emphasize practical workflows and agentic tasks beyond static academic benchmarks.
- Safety: The reported safety evaluation includes correct identity recognition and clarification of a lack of affiliation.This example concerns the relationship between DeepSeek and Doubao.
- Conclusion: Seed1.8 combines strong base LLM and VLM capabilities with multimodal perception, tool use, and multi-step task execution.The design also considers practical deployment constraints.
- Conclusion: Model development extends evaluation beyond static academic benchmarks to real-world-oriented workflows and agentic tasks.The stated goal is to incorporate practical use cases into model assessment.
6 Contributions
The supplied Contributions passages consist of an author list rather than substantive descriptions of Seed1.8’s contributions. They identify contributors by name and state how the list is organized.
- Contributors: The section provides a long list of named contributors.The supplied passages contain contributor names rather than contribution descriptions.
- Contributors: The contributor names span entries beginning with A through C in the supplied text.These passages include names from Anqi Dai through Chenyuan Wang and related entries.
- Contributors: The list continues through entries beginning with D through Z in the supplied text.The supplied passages include names from Cunwei Jie through Zuo Wang and Zuquan Song.
A The Seed Evaluation System
The Seed Evaluation System is designed to connect benchmark scores with real-world utility by prioritizing user experience, realistic scenarios, and advanced intelligence. It combines use-case analysis with application-oriented and reasoning-focused evaluations, while excluding internal health results because of data-privacy constraints.
- Evaluation Principles: The evaluation philosophy prioritizes user experience, real-world scenarios, and pushing the frontier of intelligence.The system is intended to bridge model capabilities and real-world utility rather than rely solely on synthetic tasks.
- User-Centered Evaluation: ChatGPT use-case distributions identify information seeking, text editing, and tutoring as the top three categories used to shape benchmark coverage.These observations are combined with standard agentic-LLM benchmarks to align evaluation with C-end user needs.
- Scope Boundary: Internal health-benchmark results are excluded from the report because of data-privacy constraints.This scope boundary is stated in a footnote to the use-case evaluation discussion.
- Real-World Scenarios: The system designs economically valuable tasks that mirror real-world complexity so evaluation improvements correspond to tangible usage value.These tasks are organized under the Economically Valuable Fields category.
- Intelligence Frontier: Advanced benchmarks in reasoning, mathematics, and coding measure upper-limit performance alongside usability-focused evaluation.The stated purpose is to ensure that practical-use emphasis does not compromise core intelligence.
B Details of In-house Benchmarks
This section introduces the details of the in-house benchmarks used in the paper.
- The section presents Seed1.8’s in-house benchmarks.
- These benchmarks are introduced as part of the paper’s evaluation materials.
- The section provides benchmark details for subsequent evaluation analysis.
B.1 Agentic Tasks
The paper develops agentic benchmarks covering search, coding, instruction following, dialogue, mathematical reasoning, knowledge, and artifact generation. These benchmarks target realistic, long-horizon, expert, and multimodal workflows.
- Agentic benchmark suite: New open-source benchmarks extend agentic evaluation beyond BrowseComp, SWE-bench verified, and τ 2-Bench to search and coding.
- Search and retrieval: MM-BrowseComp evaluates long-context reasoning and tool-based retrieval over text, images, and videos in simulated web browsing.
- Search and retrieval: WideSearch tests whether agents can consistently gather large amounts of scattered information accurately throughout long, repetitive tasks.
- Coding and execution: AInstein-SWE-Bench evaluates research-level scientific coding through domain understanding, repository navigation, and test-driven execution in containers.
- Coding and execution: Multi-SWE-bench measures issue resolution across seven programming languages using 1,632 instances reviewed by 68 experts from 2,456 candidates.
- Broader capability coverage: The suite also covers authentic interactive artifacts, discourse-level translation, counter-intuitive instructions, long multi-turn dialogue, advanced mathematics, graduate knowledge, biology, and practical expert knowledge.Examples include U-Artifacts, DiscoX, Inverse IFEval, MARS-Bench, BeyondAIME, SuperGPQA, BIOBench, and LPFQA.
B.4 VLM Tasks
The VLM and application-aligned evaluations span vision-grounded reasoning, education, customer support, information processing, intention recognition, structured extraction, and complex workflows.
- VLM tasks: MME-CC evaluates vision-grounded spatial, geometric, and visual-knowledge reasoning across 11 task types and 1,173 expert-annotated questions.
- Application-aligned tasks: Nine economically significant benchmarks cover six base-LLM tasks and three agentic tasks aligned with practical utility.
- Application-aligned tasks: The education benchmark tests teaching scenarios including problem solving, grading, explanation, and question generation across K–12 subjects.
- Application-aligned tasks: Customer-support tasks require product-grounded recommendations, fallback responses when information is insufficient, and preservation of original links.
- Application-aligned tasks: Information-processing tasks organize time-bounded emails by count, category, summaries, and date-range formatting.
- Application-aligned tasks: Intention recognition and structured extraction evaluate JSON intent routing and faithful EIA field extraction with units and concise formatting.
C.1 Travel Planning Assistance
The travel-planning example presents a detailed Berlin itinerary combining hotels, transport, attractions, restaurants, times, costs, and reference images.
- Itinerary: The response begins with an 00:00–08:00 hotel stay at the InterContinental Berlin before traveling to Museum für Naturkunde.
- Itinerary: The itinerary sequences museum, restaurant, and Berliner Fernsehturm visits with scheduled taxi transfers between locations.
- Itinerary: The plan includes evening dining at Rutz Restaurant followed by transport back to the InterContinental Berlin.
C.2 Expert-Level Tasks
Seed1.8’s expert-level task responses span legal analysis, financial reporting, and humanities dialogue. The responses provide case-specific conclusions, market data, and historically styled philosophical discussion.
- Legal task: The legal response concludes that Zhang’s personal unlimited joint and several liability guarantee should be deemed invalid.It attributes the conclusion to limited civil capacity at signing and the absence of legal-representative ratification.
- Legal task: Because Zhang was not at fault for the invalid guarantee, the response concludes that he need not bear compensation liability.It also identifies evidence areas relevant to the bank’s counsel, including Zhang’s civil capacity and the bank’s examination duty.
- Financial task: The financial response ranks China’s top single-country export markets for January–September 2025 and reports export values, year-on-year changes, and shares.The United States is listed first at RMB 22.77 trillion, down 16.2% year on year and accounting for 11.42% of total exports.
- Humanities task: The humanities response presents Confucius and Socrates debating life, death, moral responsibility, justice, and preserving oneself to pursue one’s mission.Their positions are expressed through dialogue grounded in the supplied historical and philosophical framing.
C.3 Scientific Research Tasks
Seed1.8’s scientific research task response analyzes optogenetic pyroptosis constructs, fluorescence labels, and treatment-dependent caspase activation. It identifies truncations that retain activity, the least effective modification, and the strongest treatment condition.
- Biology research task: The response concludes that the CARD domain can be omitted because constructs 51-435, 90-435, and 130-435 retain high pyroptosis-inducing activity under blue light.The CARD domain is identified as amino acids 2-92.
- Biology research task: The 2-435 construct is identified as the least effective modification because it has the lowest percentage of pyroptotic cells under light-on conditions.The construct retains the full CARD domain and performs worse than the other listed constructs at each time point.
- Biology research task: Fluorescent reagent A appears green, while Annexin V fluorescence appears blue in the image.The response associates reagent A with nuclei of dying cells and Annexin V with externalized phosphatidylserine.
- Biology research task: 30 minutes of blue light produces the strongest caspase-4/5 cleavage and the highest amount of cleaved GSDMD p31.The response contrasts this condition with weaker cleavage after LPS treatment and 10 minutes of blue light.
C.4 Scientific Software Engineering Tasks
The scientific software engineering section summarizes Seed1.8’s full agent response in Table 11 rather than reproducing the entire interaction. The table is identified as a structured summary of that task response.
- Scientific software engineering task: The section states that the full agent track is too long to present in full.Its response is therefore summarized in Table 11.
- Scientific software engineering task: Table 11 provides a structured summary of Seed1.8’s full response to the scientific software engineering task.The caption identifies the task as belonging to Section 3.4.
- Scientific software engineering task: The table is presented as the reporting format for the scientific software engineering task response.Its caption explicitly links Table 11 to Seed1.8’s full response.