Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

22,081 to 22,140 of 61,116

  1. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

    Yutao Mou, Pengfei Yang, Zhe Yin +6

    cs.CRcs.CLarXiv:2608.11878v12026
  2. AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

    Ivan Moshkov, Darragh Hanley, Ivan Sorokin +5

    cs.AIcs.CLcs.LGarXiv:2504.16891v12025
  3. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

    Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson +4

    cs.AIcs.CLcs.LGarXiv:2502.01100v22025
  4. Dion3: Full-Stack Orthogonal Updates

    Noah Amsel, Jack Zhang, Kwangjun Ahn +5

    cs.LGcs.AIarXiv:2608.11612v12026
  5. FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

    Yuren Cong, Mengmeng Xu, Christian Simon +7

    cs.CVarXiv:2310.05922v32023
  6. Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution

    Zhiyuan You, Xin Cai, Jinjin Gu +2

    cs.CVarXiv:2501.11561v32025
  7. From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

    Zepeng Wang, Jiagao Hu, Fuhao Li +3

    cs.CVcs.AIeess.IVarXiv:2608.11562v12026
  8. Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

    Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae +1

    cs.CVarXiv:2608.10708v12026
  9. InSight-doc: Agentic Visual Perception for Long-Document Understanding

    Kaican Li, Weiyan Xie, Lewei Yao +4

    cs.CVcs.CLcs.LGarXiv:2608.10628v12026
  10. A deep learning model for estimating story points

    Morakot Choetkiertikul, Hoa Khanh Dam, Truyen Tran +3

    cs.SEcs.LGstat.MLarXiv:1609.00489v22016
  11. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

    Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3

    cs.AIcs.CLarXiv:2608.10366v12026
  12. SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

    Yuling Shi, Jinghan Xu, Kelin Fu +12

    cs.CLcs.SEarXiv:2608.09802v12026
  13. Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks

    Ziwei Ji, Matus Telgarsky

    cs.LGmath.OCstat.MLarXiv:1909.12292v42019
  14. Evo-Bench: Can Language Models Improve Agent Harness?

    Lisheng Huang, Chen Yang, Hao Zhou +6

    cs.CLarXiv:2608.09096v22026
  15. Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

    Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20

    cs.CLcs.AIarXiv:2608.09766v12026
  16. Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning

    Kimin Lee, Kibok Lee, Jinwoo Shin +1

    cs.LGstat.MLarXiv:1910.05396v32019
  17. Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

    Víctor Gallego

    cs.LGcs.AIarXiv:2608.08722v12026
    Summaries:한국어
  18. Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

    Harshitha Kolukuluru, Reshma Ashok, Kirat Arora +7

    cs.AIcs.IRcs.MAarXiv:2608.08389v12026
  19. CEM-RL: Combining evolutionary and gradient-based methods for policy search

    Aloïs Pourchot, Olivier Sigaud

    cs.LGcs.NEstat.MLarXiv:1810.01222v32018
  20. A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization

    Víctor Gallego

    cs.LGcs.NEarXiv:2608.08156v12026
  21. Thought-Level Beam Search for Reasoning

    Lijie Yang, Hongyin Luo, Jiawei Zhao +2

    cs.AIarXiv:2608.08020v22026
  22. Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

    Changzhi Liu, Yilun Liu, Sikuan Yan +2

    cs.AIcs.LGarXiv:2608.07645v12026
  23. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    Tao Feng, Fangxu Yu, Haozhen Zhang +9

    cs.CLarXiv:2608.06867v12026
  24. Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

    Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran

    math.OCcs.LGarXiv:2608.07265v22026
  25. End-to-End Lane Marker Detection via Row-wise Classification

    Seungwoo Yoo, Heeseok Lee, Heesoo Myeong +4

    cs.CVcs.LGarXiv:2005.08630v12020
  26. DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text

    Xianjun Yang, Wei Cheng, Yue Wu +3

    cs.CLcs.AIarXiv:2305.17359v22023
  27. RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

    Shitao Xiao, Zheng Liu, Yingxia Shao +1

    cs.CLarXiv:2205.12035v22022
  28. General-Reasoner: Advancing LLM Reasoning Across All Domains

    Xueguang Ma, Qian Liu, Dongfu Jiang +3

    cs.CLarXiv:2505.14652v52025
  29. Small Foundation Models of Human Cognition and Behaviour

    Nick Oh, Fernand Gobet

    cs.AIcs.CYarXiv:2608.05224v32026
  30. Delving into LLM-assisted writing in biomedical publications through excess vocabulary

    Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát +1

    cs.CLcs.AIcs.CYarXiv:2406.07016v52024
  31. StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

    Meng Wei, Chenyang Wan, Xiqian Yu +9

    cs.ROcs.CVarXiv:2507.05240v22025
  32. b-Bit Minwise Hashing

    Ping Li, Arnd Christian Konig

    cs.DScs.DBcs.IRarXiv:0910.3349v12009
  33. SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models

    Yuzhou Huang, Liangbin Xie, Xintao Wang +8

    cs.CVarXiv:2312.06739v12023
  34. SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

    Yue Zhang, Yingzhao Jian, Yunqiu Xu +2

    cs.CVarXiv:2608.05137v32026
  35. MatrAIx: Simulating the World with 8.3 Billion Persona Agents

    Xiaomin Li, Yuexing Hao, Jianheng Hou +90

    cs.AIarXiv:2608.04205v12026
  36. DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

    Boyan Li, Zhuowen Liang, Yupeng Xie +11

    cs.AIarXiv:2608.03451v12026
  37. Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

    Ayoub Kirouane, Christos Petrocheilos

    eess.AScs.AIcs.CLarXiv:2608.05138v12026
  38. When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

    Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague +2

    cs.AIstat.MLarXiv:2608.03506v12026
  39. LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

    Ziyu Ma, Hailang Huang, Shun Zou +5

    cs.CVarXiv:2608.01964v12026
  40. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

    Zixuan Huang, Yang Zhou, Kaixuan Wang +7

    cs.AIarXiv:2608.01755v22026
  41. On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

    Yongliang Wu, Yizhou Zhou, Zhou Ziheng +7

    cs.LGarXiv:2508.05629v32025
  42. VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

    Jiaxin Bai, Jiaxuan Xiong

    cs.CLcs.ROarXiv:2607.25236v12026
  43. When Tokenization is Secretly Output Supervision

    Tanja Baeumel, Josef van Genabith, Simon Ostermann

    cs.CLarXiv:2609.01386v12026
  44. Spectral Prior for Reducing Exposure Bias in Diffusion Models

    Yuya Kobayashi, Masato Ishii, Yuhta Takida +2

    cs.CVarXiv:2607.22091v12026
  45. FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

    Xianfu Cheng, Shiwei Zhang, Jiyu Zhao +10

    cs.CEarXiv:2607.19238v12026
  46. Recursive Harness Self-Improvement

    Hyunin Lee, Jinglue Xu, Jeffrey Seely +3

    cs.LGcs.AIarXiv:2607.15524v12026
    Summaries:한국어
  47. The Horseshoe+ Estimator of Ultra-Sparse Signals

    Anindya Bhadra, Jyotishka Datta, Nicholas G. Polson +1

    math.STarXiv:1502.00560v22015
  48. OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

    Mengkang Hu, Yuhang Zhou, Wendong Fan +13

    cs.AIcs.CLarXiv:2505.23885v22025
  49. ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

    Xiaozhong Lyu, Gen Li, Zhiyin Qian +3

    cs.CVcs.AIarXiv:2607.17790v12026
  50. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

    Ruhan Wang, Yucheng Shi, Zongxia Li +7

    cs.AIcs.SEarXiv:2607.13285v12026
  51. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

    Qingyu Zhang, Qianhao Yuan, Hongyu Lin +5

    cs.LGcs.AIcs.CLarXiv:2607.13124v22026
    Summaries:한국어
  52. Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

    Yubo Wang, Jiarong Liang, Yuxuan Zhang +5

    cs.AIcs.CLarXiv:2607.12463v32026
  53. The AI Index 2021 Annual Report

    Daniel Zhang, Saurabh Mishra, Erik Brynjolfsson +10

    cs.AIcs.GLarXiv:2103.06312v12021
  54. Classification and Clustering of Arguments with Contextualized Word Embeddings

    Nils Reimers, Benjamin Schiller, Tilman Beck +3

    cs.CLarXiv:1906.09821v12019
  55. 2018 Robotic Scene Segmentation Challenge

    Max Allan, Satoshi Kondo, Sebastian Bodenstedt +38

    cs.CVcs.ROarXiv:2001.11190v32020
  56. Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents

    Ruochen Zhou, Zhengyu Chen, Luan Zhang +3

    cs.CLarXiv:2609.01294v12026
  57. Unified Reward Model for Multimodal Understanding and Generation

    Yibin Wang, Yuhang Zang, Hao Li +2

    cs.CVarXiv:2503.05236v22025
  58. LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs

    Muhammed Saeed, Simon Razniewski

    cs.CLarXiv:2609.01182v12026
  59. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

    Chen Tang, Yizhou Wang, Jianyu Wu +26

    cs.CLcs.AIcs.CEarXiv:2607.07708v12026
    Summaries:한국어
  60. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

    Andrey Podivilov, Vadim Lomshakov, Sergey Savin +4

    cs.AIcs.LGcs.SEarXiv:2607.06624v22026