Source-linked AI summary
From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought
Lionel Wong, Gabriel Grand, Alexander K. Lew, Noah D. Goodman, Vikash K. Mansinghka, Jacob Andreas, Joshua B. Tenenbaum
TL;DR
The paper asks how language can robustly inform general cognition despite gaps in current and symbolic systems. It proposes rational meaning construction, showing language-to-probabilistic-program translations support reasoning and world-model construction.
Problem
The paper asks how language gets meaning and can robustly support systems modeling the full scope of human thought, beyond brittle symbolic mappings.
Method
Rational meaning construction combines probabilistic programs for structured situations and belief updating with language models translating utterances into probabilistic language-of-thought expressions.
Results
Across probabilistic, relational, physical, and social examples, the framework integrates language with reasoning, including probability distributions from simulated generative models.
Takeaways & Limitations
Language can support autonomous construction of new concepts and world models that enable coherent downstream reasoning.
Takeaways & Limitations
The kinship model simplifies relationships for tractable inference and can encode gender biases through binary kinship terms and gender-associated names.
Abstract
from arXiv · showhide
How does language inform our downstream thinking? In particular, how do humans make meaning from language--and how can we leverage a theory of linguistic meaning to build machines that think in more human-like ways? In this paper, we propose rational meaning construction, a computational framework for language-informed thinking that combines neural language models with probabilistic models for rational inference. We frame linguistic meaning as a context-sensitive mapping from natural language into a probabilistic language of thought (PLoT)--a general-purpose symbolic substrate for generative world modeling. Our architecture integrates two computational tools that have not previously come together: we model thinking with probabilistic programs, an expressive representation for commonsense reasoning; and we model meaning construction with large language models (LLMs), which support broad-coverage translation from natural language utterances to code expressions in a probabilistic programming language. We illustrate our framework through examples covering four core domains from cognitive science: probabilistic reasoning, logical and relational reasoning, visual and physical reasoning, and social reasoning. In each, we show that LLMs can generate context-sensitive translations that capture pragmatically-appropriate linguistic meanings, while Bayesian inference with the generated programs supports coherent and robust commonsense reasoning. We extend our framework to integrate cognitively-motivated symbolic modules (physics simulators, graphics engines, and planning algorithms) to provide a unified commonsense thinking interface from language. Finally, we explore how language can drive the construction of world models themselves. We hope this work will provide a roadmap towards cognitive models and AI systems that synthesize the insights of both modern and classical computational perspectives.
1 Introduction
The paper asks how language can inform general cognition and proposes rational meaning construction as a bridge between broad-coverage language models and structured symbolic world modeling. It presents a prospective architecture in which language is translated into probabilistic programs that support coherent inference and can grow new concepts and world models.
- Human language supports communication about beliefs, uncertainty, perception, hypothetical futures, goals, plans, explanations, and theories.
- Traditional symbolic models support structured world modeling and inference but face persistent scalability, scope, and language-to-symbol mapping challenges.Their domain representations and semantic parsers were often bespoke, hand-engineered, or learned with strong supervision.
- Large language models provide broad linguistic coverage but can fail on minor input changes and produce confident language inconsistent with calibrated truth or belief.
- Rational meaning construction integrates a probabilistic language of thought for structured world models with a mechanism that maps natural-language utterances to distributions over probabilistic-language-of-thought expressions.
- The framework uses language to guide, constrain, and drive downstream world modeling and inference, including the autonomous construction of new concepts and whole world models without hand engineering.
- The paper is prospective: its examples motivate future work on language acquisition and production, scalable inference, robust translation, and learning.
2 Overview of the key ideas
Rational meaning construction links language to thought by translating natural-language utterances into probabilistic language-of-thought expressions, where inference produces context-sensitive beliefs about possible worlds. The framework combines LLM-based broad-coverage translation with probabilistic programs and illustrates how this supports flexible, interpretable reasoning from language.
- 2.1 Our proposal: A framework for modeling rational meaning construction: The framework proposes rational meaning construction, using LLMs to map language into probabilistic language-of-thought expressions that support reasoning about possible worlds.The mapping is contextual and distributional rather than a single fixed interpretation.
- 2.1.3 Illustrating the architecture by example: The examples are illustrative and pedagogical rather than comprehensive, with each domain requiring further work to scale toward richer models of language and cognition.The authors present the examples as representative starting points for broader computational cognitive models and more systematic evaluation.
- 2.2 Understanding language with probabilistic reasoning: LLMs serve as broad-coverage meaning functions because language-code pretraining and prompting support translation across varied phrasing while respecting the probabilistic program’s domain-specific structure.The approach combines invariance to superficial linguistic variation with sensitivity to wording that affects meaning.
- 2.2 Understanding language with probabilistic reasoning: Probabilistic inference performed by the probabilistic language of thought yields richer and more interpretable beliefs than textual responses alone, including posterior distributions over latent traits and match outcomes.In the tug-of-war example, the model estimates Josh’s strength and Gabe’s chance of winning against Josh at 23.90%.
- 2.2 Understanding language with probabilistic reasoning: The framework updates beliefs flexibly when new or unlikely information changes the interpretation of observed events, shifting inferred strength according to competing hypotheses about Josh’s laziness.Gabe’s win implies different strength estimates depending on whether Josh is likely or unlikely to have thrown the match.
- 2.2 Understanding language with probabilistic reasoning: In a more complex dialogue, appropriately prompted LLMs translate observations and questions into condition and query statements that support defeasible, flexible inferences about teams and match outcomes.The framework also allows newly defined concepts, such as faculty-team and student-team, to be referenced symbolically in later reasoning.
3 Understanding and reasoning about language with world models
The framework applies language-to-program translation and probabilistic world modeling to logical, relational, visual, physical, and social reasoning. Kinship provides a compact testbed where natural-language statements constrain possible family worlds and support mixed deductive–inductive inference.
- The framework extends language-informed world modeling across logical, relational, visual, physical, and social domains.Each domain uses generative programs to interpret language as observations or questions about uncertain situations.
- Kinship as a reasoning domain: Kinship is a useful testbed because its language expresses composition, symmetry, ambiguity, and both deductive and inductive relations.Examples range from “my mother’s father” to ambiguous references such as “Blake’s uncle.”
- Generative kinship model: The kinship domain theory generates family genealogies through recursive random choices about births, partnerships, and children.Each generated family tree represents a possible world whose relationships can be queried.
- Scope and limitations: The toy kinship model omits important identity and relationship nuances, while gendered English terms can push its predicates toward traditional gender assignments.The authors identify cultural and social extensions as opportunities for future development.
- Conceptual representation: A conceptual system represents kinship relations with derived predicates built over low-level tree-accessor functions.These predicates provide a higher-level interface for expressing relational statements about individuals.
- Language-to-program translation: The language model translates natural-language kinship statements into conditioning and query programs using the domain model and predicate definitions.The resulting interface connects flexible, context-aware language grounding with explicit family-tree reasoning.
- Inference from language: Each translated utterance constrains family-tree samples, reducing uncertainty and producing posterior hypotheses about the family described.Successive conditions progressively crystallize the model’s picture of the relevant family tree.
A. Language-to-code translation B. Family trees sampled from conditioned kinship domain theory
The framework translates language into probabilistic constraints and queries over structured world models, supporting kinship inference and language-grounded visual scene reasoning. Conditioned samples represent possible worlds, while rendering and inference connect linguistic descriptions to imagined scenes.
- A. Language-to-code translation: Kinship utterances become Church conditions that cumulatively constrain sampled family trees, enabling posterior inference over unknown relations.The conditioned samples represent possible family trees consistent with the growing set of linguistic constraints.
- B. Family trees sampled from conditioned kinship domain theory: Contradictory information can reduce a previously likely kinship answer to 0%, showing principled updating over alternative family configurations.After conditioning on Blake having two children, Blake is assigned > 80% probability as Dana’s parent; the probability drops to 0% when Dana is specified as an only child.
- A. Language-to-code translation: The visual extension combines probabilistic scene models with a graphics renderer so language can specify conditions and queries over imagined tabletop scenes.The scene domain includes colored mugs, cans, and bowls, with symbolic world states rendered into visual depictions.
- B. Family trees sampled from conditioned kinship domain theory: The scene implementation is illustrative, while richer visual reasoning would need factors such as lighting, viewpoint, stereo depth, and perceptual noise.These factors are identified as examples modeled in related scene-understanding work rather than fully represented by this implementation.
- A. Language-to-code translation: Translations generalize across conjunctions, syntactic variation, negation, quantities, object-property compositions, and comparative relations.The examples include expressions such as “there aren’t any” and comparisons between red mugs and green cans.
A. Language-to-code translation B. Rendered scenes from conditioned generative model
The framework translates language into composable probabilistic program expressions that condition and query generative world models. These models connect linguistic descriptions to rendered visual scenes, physical simulations, and probabilistic reasoning about agents and plans.
- A. Language-to-code translation: LLMs translate syntactic variation, conjunctions, negation, comparatives, and vague expressions into composable program conditions.The translations include context-sensitive thresholds and interpretations over object sets.
- B. Rendered scenes from conditioned generative model: Sampling from conditioned scene distributions and rendering the symbolic states produces images consistent with successive linguistic descriptions.Each sentence updates the underlying distribution, allowing increasingly constrained scene interpretations.
- B. Rendered scenes from conditioned generative model: Semantic event predicates such as hitting, moving, and resting are derived from continuous physical states in the integrated physics engine.This provides an interface between probabilistic programs and continuous physical simulation.
- B. Rendered scenes from conditioned generative model: Physics interfaces turn language about mass, shape, friction, and force into simulated scene distributions that support downstream event inferences.Queries over these distributions are answered by sampling and running physical simulations over possible world states.
- B. Rendered scenes from conditioned generative model: The physical-scene implementation remains a deliberately simple world model, omitting richer environments, materials, object configurations, and arbitrary forces.The authors present these extensions as future directions rather than implemented capabilities.
- B. Rendered scenes from conditioned generative model: The framework does not implement a perceptual module, though the authors describe future integration with inverse graphics and visual observations.Such integration is proposed for joint reasoning about visible scenes and latent physical properties.
- A. Language-to-code translation: Model-based planners support forward and inverse inferences about agents by unifying goals, actions, observations, and latent world states.The framework can infer plans and update expectations about agents from new observations, including indirect evidence about restaurant availability.
- A. Language-to-code translation: In social reasoning, preference terms map to context-specific utility thresholds, and quantifiers map to conjunctions over contextual restaurant sets.The LLM generalizes from example translations to additional preference terms such as hate and love.
4 Growing and constructing world models from language
The paper extends its language-to-code framework from interpreting language within fixed domains to learning new concepts and constructing probabilistic world models. Language can define program expressions that enrich existing models or assemble new models from domain descriptions.
- 4 Growing and constructing world models from language: The paper asks how world models expressed as probabilistic programs can be learned rather than hand-coded.This is framed as a scalability question for the PLoT account of language understanding.
- 4 Growing and constructing world models from language: The framework is presented as a way to learn concepts and domains in language that can support future observations, queries, relational reasoning, goals, and plans.This consequence is stated as a near-term direction rather than a completed general solution.
- 4.1 Growing a world model from language: The extended kinship model can use newly defined relations productively in subsequent sentences and probabilistic reasoning.The examples include culturally and linguistically distinct kinship concepts such as pibling and Northern Paiute pāan’i.
- 4.1 Growing a world model from language: A language-to-code LLM converts linguistic definitions of new kinship relations into program expressions that extend an existing generative model.The approach supports concepts from contemporary English and Northern Paiute.
- 4.1 Growing a world model from language: Language can specify new parts of a world model as program expressions, which then become a structured basis for reasoning about observations.The same principle is proposed for extending domains beyond kinship.
- 4.2 Constructing new world models from language: The proposed world-model construction view would treat language as conditioning uncertainty over possible model fragments, but the paper instead directly samples fragments consistent with each statement.The shortcut avoids the harder problem of Bayesian inference over complete world models.
A. Prompt, containing unrelated example world model
The paper demonstrates constructing a probabilistic tug-of-war world model from language using unrelated example code and line-by-line translation. The resulting model recovers the essential structure of a hand-coded model, while the broader construction examples remain limited in scope.
- A. Prompt, containing unrelated example world model: Codex constructs a tug-of-war generative model from language, using unrelated medical-diagnosis code as an example world model.The model is generated line by line from linguistic explanations of the target domain.
- A. Prompt, containing unrelated example world model: The generated model is semantically equivalent to the earlier hand-coded model despite superficial differences in naming and parameter choices.The specific priors may vary slightly, but the essential structure is recovered.
- A. Prompt, containing unrelated example world model: Once constructed, the new domain model supports interpreting observations, answering queries, and adding further definitions.The paper illustrates this with future observations and questions about strengths in the tug-of-war domain.
- A. Prompt, containing unrelated example world model: The construction examples are limited to cases with an explicit connection between linguistic instructions and probabilistic-program expressions.The authors note that real language often provides indirect clues or is absent when world models are assembled from experience.
5 Open questions and future directions
The paper identifies open challenges in scaling probabilistic inference, integrating structured models with neural language systems, and connecting language-informed reasoning to cognitive and neural evidence. It also points toward flexible hybrid architectures, improved robustness and interpretability, and more data-efficient language learning.
- Inference and world modeling: A major remaining challenge is building world-modeling modules that fit sparse data, simulate rare events, and plan under uncertainty.The paper notes that existing game engines lack these affordances, while neurally guided program learning remains incomplete.
- Inference and world modeling: Scaling probabilistic inference to human-like robustness, speed, efficiency, and flexibility remains an explicitly unaddressed challenge.Rejection sampling requires exponentially more proposal attempts as scenarios become less likely under the prior.
- Structured hybrid models: Future systems could flexibly allocate language understanding between explicit probabilistic inference and learned amortized prediction, trading speed against accuracy.The proposed research question is how to identify computations that can be reliably emulated by learned components.
- Empirical directions: Preliminary applications span physical commonsense reasoning, social reasoning about goal-directed agents, and human-like scalar-implicature judgments.These examples are presented as promising preliminary results rather than a complete evaluation.
- Language and thought in the brain: Neuroscientific evidence motivates viewing LLMs as context-aware mappings from language to meanings rather than end-to-end models of language and reasoning.The framework also suggests that code-trained LLMs may better capture latent semantic and syntactic structure than language-only LLMs.
- Structured hybrid models: The framework treats a probabilistic language of thought as a unifying substrate that can express world models and nest cognitively motivated modules.This reverses the view that probabilistic programs are merely plug-ins for language models.
6 Conclusion
The conclusion argues that human-like language understanding requires models grounded in purposeful world models and beliefs rather than language statistics alone. Such a theory could support AI systems that understand language reliably while remaining interpretable, explainable, and controllable.
- Conclusion: Human language understanding is grounded in world models and beliefs constructed toward intentions and desires, unlike the language-learning regime of current LLMs.The paper contrasts human learning from far less input with models trained on massive language datasets.
- Conclusion: A cognitive theory of language must explain how people generate novel thoughts, language, words, and even languages.The conclusion presents this as part of capturing the human relationship between language and thought.
- Conclusion: The proposed direction could form the basis for AI models that understand people reliably and predictably while remaining interpretable, explainable, and controllable.The passage presents this as a possible consequence of a cognitive theory of human language.
Appendices
The appendices provide code for interpreting the paper’s examples and direct readers to the repository for the most complete and reproducible implementation.
- Appendices: Reference code is included with human-readable comments to help interpret the paper’s examples.The appendix cautions that this code may not be the most up-to-date version.
- Appendices: The GitHub repository contains the most complete and corrected code for the examples, along with execution and reproducibility instructions.The repository is identified as github.com/gabegrand/world-models.
A.1 Probabilistic reasoning
The probabilistic-reasoning appendix encodes a tug-of-war as a generative probabilistic program, then translates natural-language conditions and queries into executable program expressions.
- Generative domain theory: The tug-of-war model assigns each player a Gaussian-distributed strength centered at 50 with standard deviation 20.Strength is represented as a memoized function over players.
- Generative domain theory: A team’s strength sums its players’ contributions, with lazy players pulling at half their individual strength.The team-strength function applies this adjustment player by player.
- Generative domain theory: The winner is defined as the team with greater total strength, and won-against returns a Boolean outcome.The model compares the two teams’ computed strengths.
- Prompt interface: Condition statements provide scenario facts, whereas Query statements evaluate quantities of interest.The examples include Alice winning against Bob as a condition and asking whether Sue is stronger than Mary as a query.
- Prompt interface: The translation interface handles underspecified language by converting “Sue is very strong” into a strength threshold above 75.It also supports defining reusable predicates such as stronger-than? for subsequent conditions and queries.
- Prompt interface: The examples express plural relations and hypothetical questions through programmatic conditions and queries over player groups.Examples include at least two players stronger than John and asking who would win if Mary played Tom.
A.2.1 Generative world model for kinship
The kinship world model generates probabilistic family trees, assigns people attributes and names, and exposes utilities for querying kinship relations. Conditions and queries translate natural-language statements into structured constraints and questions over the generated tree.
- Generative world model: People receive unique identifiers, randomly assigned genders, shuffled names, and optional names added across the generated tree.Named people are paired with entries from the conversational name list, while unnamed people remain addressable by person identifier.
- Generative world model: The generative model creates a family tree recursively from parent nodes, probabilistic partnerships, bounded child counts, and depth limits.Each generated person stores an identifier, name-related information, gender, and parent identifiers; child identifiers are linked back to both parents.
- Kinship utilities: Kinship utilities retrieve people by name or identifier, access properties, filter trees, and test predicates over family relationships.The utilities support parents, grandparents, children, siblings, partners, and gender-specific relations such as father, mother, brother, and sister.
- Kinship utilities: The conceptual system derives relations such as parents, grandparents, children, siblings, partners, and gender-specific kinship from stored tree properties.Derived predicates compose simpler lookup and membership operations to express relational queries over the family tree.
- Translation examples: Condition statements provide scenario facts, whereas query statements evaluate quantities or relationships of interest.Examples include asserting family relations, checking sibling counts, finding parents or grandparents, counting children, and testing whether someone has a sister.
A.2.5 Why not Prolog?
The paper uses Church rather than Prolog because probabilistic programs support mixed deductive and inductive reasoning over uncertain world models. Prolog remains compatible with some translations, but its definite-clause semantics constrain expressible inferences.
- Alternative languages: LLMs can translate natural-language kinship utterances into Prolog, and the semantic-parsing interface is not tied to Church.The paper notes that other targets, including Prolog, SMT solvers, and Python, could replace Church with suitable prompting.
- Limitations of Prolog: Prolog’s unidirectional definite-clause implication makes reverse and abductive consequences of facts difficult to represent straightforwardly.A stated grandfather relation does not naturally provide the bidirectional or explanatory inferences that the relation’s rule might suggest.
- Limitations of Prolog: Prolog also poorly captures inductive inferences such as higher likelihoods of additional children or marriage given that someone has a child.These inferences require uncertainty over the world model rather than only logical entailment.
- Motivation: Church is chosen for kinship because the framework needs to incorporate and trade off uncertainty alongside deductive reasoning.The paper contrasts this with probabilistic Prolog extensions and emphasizes Church’s broader probabilistic-programming approach.
- Visual-domain examples: The visual examples represent scenes as sets of objects with attributes such as shape and color, then translate natural-language conditions and queries into filters and counts.The examples include checking for colored objects, counting objects of a shape, and combining multiple scene constraints.
A.3.4 Translation examples for visual scenes
The translations express visual and physical-scene descriptions as probabilistic-program conditions over object properties, events, and queried outcomes. Related examples also encode gridworld actions, utilities, transitions, and policies for social reasoning.
- Relational events: Relational event translations compose predicates for multiple objects and an event, including a red cube hitting a blue cube.The program identifies event subject and object roles and labels the event as hitting.
- Physical queries: The physical-scene examples support downstream queries by translating questions about event attributes, such as a red block’s final velocity.The query retrieves the subject’s final velocity from the represented event.