Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,221 to 5,280 of 18,866

  1. NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

    NextStep Team, Chunrui Han, Guopeng Li +47

    cs.CVarXiv:2508.10711v22025
  2. Asymmetric Bilateral Motion Estimation for Video Frame Interpolation

    Junheum Park, Chul Lee, Chang-Su Kim

    cs.CVarXiv:2108.06815v12021
  3. FLAVR: Flow-Agnostic Video Representations for Fast Frame Interpolation

    Tarun Kalluri, Deepak Pathak, Manmohan Chandraker +1

    cs.CVarXiv:2012.08512v32020
  4. Multi-Modal Answer Validation for Knowledge-Based VQA

    Jialin Wu, Jiasen Lu, Ashish Sabharwal +1

    cs.CVcs.CLarXiv:2103.12248v32021
  5. On the Effectiveness of Image Rotation for Open Set Domain Adaptation

    Silvia Bucci, Mohammad Reza Loghmani, Tatiana Tommasi

    cs.CVarXiv:2007.12360v12020
  6. WoW: Towards a World omniscient World model Through Embodied Interaction

    Xiaowei Chi, Peidong Jia, Chun-Kai Fan +33

    cs.ROcs.CVcs.MMarXiv:2509.22642v22025
  7. Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

    Mehrdad Fazli, Sina Mansouri, Mohit Marvania +1

    cs.CVarXiv:2609.01888v12026
  8. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    Yongliang Wu, Zonghui Li, Xinting Hu +7

    cs.CVarXiv:2505.16707v12025
  9. Ola: Pushing the Frontiers of Omni-Modal Language Model

    Zuyan Liu, Yuhao Dong, Jiahui Wang +4

    cs.CVcs.CLcs.MMarXiv:2502.04328v32025
  10. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

    Zhe Kong, Feng Gao, Yong Zhang +5

    cs.CVarXiv:2505.22647v12025
  11. Accuracy comparison across face recognition algorithms: Where are we on measuring race bias?

    Jacqueline G. Cavazos, P. Jonathon Phillips, Carlos D. Castillo +1

    cs.CVcs.LGarXiv:1912.07398v22019
  12. Dual Graph Convolutional Network for Semantic Segmentation

    Li Zhang, Xiangtai Li, Anurag Arnab +3

    cs.CVarXiv:1909.06121v32019
  13. LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

    Hongyu Li, Jinyu Chen, Ziyu Wei +5

    cs.CVarXiv:2501.08282v22025
  14. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

    Xingyu Fu, Minqian Liu, Zhengyuan Yang +6

    cs.CVcs.CLarXiv:2501.05452v12025
  15. A Survey of LLM-based Agents in Medicine: How far are we from Baymax?

    Wenxuan Wang, Zizhan Ma, Zheng Wang +5

    cs.CLcs.AIcs.CVarXiv:2502.11211v22025
  16. Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

    NVIDIA, :, Hassan Abu Alhaija +38

    cs.CVcs.AIcs.LGarXiv:2503.14492v22025
  17. Selective Structured State-Spaces for Long-Form Video Understanding

    Jue Wang, Wentao Zhu, Pichao Wang +4

    cs.CVarXiv:2303.14526v12023
  18. InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams

    Shuai Yuan, Yantai Yang, Xiaotian Yang +4

    cs.CVarXiv:2601.02281v12026
  19. Ming-Omni: A Unified Multimodal Model for Perception and Generation

    Inclusion AI, Biao Gong, Cheng Zou +55

    cs.AIcs.CLcs.CVarXiv:2506.09344v12025
  20. FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning

    Nam Hyeon-Woo, Moon Ye-Bin, Tae-Hyun Oh

    cs.LGcs.CVarXiv:2108.06098v32021
  21. StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems

    Jinhui Ye, Ning Gao, Senqiao Yang +7

    cs.ROcs.AIcs.CVarXiv:2604.11757v12026
  22. LLM Post-Training: A Deep Dive into Reasoning Large Language Models

    Komal Kumar, Tajamul Ashraf, Omkar Thawakar +7

    cs.CLcs.CVarXiv:2502.21321v22025
  23. A CNN-based methodology for breast cancer diagnosis using thermal images

    Juan Zuluaga-Gomez, Zeina Al Masry, Khaled Benaggoune +2

    cs.CVeess.IVarXiv:1910.13757v12019
  24. Real-Time Scene-Adaptive Tone Mapping for High-Dynamic Range Object Detection

    Gongzhe Li, Linwei Qiu, Peibei Cao +3

    cs.CVarXiv:2608.30400v12026
  25. SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

    Yiming Luo, Rongqiang Zhao, Jie Liu

    cs.LGcs.CVarXiv:2609.01051v12026
  26. Semi-Dense 3D Reconstruction with a Stereo Event Camera

    Yi Zhou, Guillermo Gallego, Henri Rebecq +3

    cs.CVcs.ROarXiv:1807.07429v12018
  27. StreamForest: Efficient Online Video Understanding with Persistent Event Memory

    Xiangyu Zeng, Kefan Qiu, Qingyu Zhang +9

    cs.CVarXiv:2509.24871v12025
  28. Automated Movie Generation via Multi-Agent CoT Planning

    Weijia Wu, Zeyu Zhu, Mike Zheng Shou

    cs.CVarXiv:2503.07314v12025
  29. DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving

    Zhenjie Yang, Yilin Chai, Xiaosong Jia +5

    cs.CVcs.AIcs.ROarXiv:2505.16278v22025
  30. Inverse Rendering for Modeling with Line Primitives

    Kenji Tojo, Ariel Shamir, Nobuyuki Umetani +1

    cs.GRcs.CVarXiv:2609.00625v12026
  31. MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation

    Zhijian Qiao, Xinjiang Wang, Jiajie Chen +5

    cs.CVcs.ROarXiv:2609.01252v12026
  32. Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

    Yunxin Li, Zhenyu Liu, Zitao Li +19

    cs.CVcs.CLarXiv:2505.04921v22025
  33. NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors

    Congyue Deng, Chiyu "Max'' Jiang, Charles R. Qi +4

    cs.CVarXiv:2212.03267v12022
  34. Wavelet Diffusion Models are fast and scalable Image Generators

    Hao Phung, Quan Dao, Anh Tran

    cs.CVeess.IVarXiv:2211.16152v22022
  35. VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control

    Yuxuan Bian, Zhaoyang Zhang, Xuan Ju +4

    cs.CVcs.AIcs.MMarXiv:2503.05639v32025
  36. Streaming Long Video Understanding with Large Language Models

    Rui Qian, Xiaoyi Dong, Pan Zhang +4

    cs.CVarXiv:2405.16009v12024
  37. Improving Object Localization with Fitness NMS and Bounded IoU Loss

    Lachlan Tychsen-Smith, Lars Petersson

    cs.CVarXiv:1711.00164v32017
  38. Hessian-based Analysis of Large Batch Training and Robustness to Adversaries

    Zhewei Yao, Amir Gholami, Qi Lei +2

    cs.CVcs.LGstat.MLarXiv:1802.08241v42018
  39. Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key

    Zhihe Yang, Xufang Luo, Dongqi Han +2

    cs.CVarXiv:2501.09695v22025
  40. MASTER: Multi-Aspect Non-local Network for Scene Text Recognition

    Ning Lu, Wenwen Yu, Xianbiao Qi +4

    cs.CVarXiv:1910.02562v32019
  41. 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

    Yipeng Gao, Lei Shu, Genzhi Ye +5

    cs.CVcs.AIcs.GRarXiv:2606.01057v12026
  42. Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

    Nuria Alina Chandra, Hannah Lee, Ryan Murtfeldt +10

    cs.CVcs.AIcs.CYarXiv:2503.02857v52025
  43. LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent

    Jianing Yang, Xuweiyi Chen, Shengyi Qian +4

    cs.CVcs.AIcs.CLarXiv:2309.12311v12023
  44. PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

    Soroush Nasiriany, Fei Xia, Wenhao Yu +20

    cs.ROcs.CLcs.CVarXiv:2402.07872v12024
  45. Pruning and Quantization for Deep Neural Network Acceleration: A Survey

    Tailin Liang, John Glossner, Lei Wang +2

    cs.CVcs.AIarXiv:2101.09671v32021
  46. Transparency of Deep Neural Networks for Medical Image Analysis: A Review of Interpretability Methods

    Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee +1

    eess.IVcs.AIcs.CVarXiv:2111.02398v12021
  47. FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery

    Xian Sun, Peijin Wang, Zhiyuan Yan +11

    cs.CVarXiv:2103.05569v22021
  48. VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction

    Hiroto Nakata, Yawen Zou, Shunsuke Sakai +5

    cs.CVarXiv:2603.13964v12026
  49. Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

    Yumi Lee, Harim Oh, Hyoryung Kim +52

    cs.CVcs.AIarXiv:2609.00866v12026
  50. Video models are zero-shot learners and reasoners

    Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6

    cs.LGcs.AIcs.CVarXiv:2509.20328v22025
  51. World Simulation with Video Foundation Models for Physical AI

    NVIDIA, :, Arslan Ali +87

    cs.CVcs.AIcs.LGarXiv:2511.00062v22025
  52. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

    Junxiang Xu, Ruisi Wang, Fanyi Pu +49

    cs.CVcs.AIcs.LGarXiv:2608.26105v12026
  53. Stitched Value Model for Diffusion Alignment

    Hyojun Go, Hyungjin Chung, Prune Truong +8

    cs.CVcs.AIcs.LGarXiv:2605.19804v12026
  54. Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

    Shuo Liang, Yixing Ma, Pengfei Zhou +32

    cs.CVcs.AIarXiv:2608.14391v12026
  55. ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

    Fan Jiang, Zhaoxu Sun, Mengchao Wang +38

    cs.CVcs.AIcs.LGarXiv:2607.19191v12026
  56. SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

    Yuyang Zhao, Yicheng Pan, Qiyuan He +6

    cs.CVcs.AIarXiv:2605.30409v12026
  57. PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation

    Yu Tian, Xintong Jiang, Jan Franklin Adamowski +2

    cs.CVarXiv:2609.02860v12026
  58. Deep Visual Domain Adaptation: A Survey

    Mei Wang, Weihong Deng

    cs.CVarXiv:1802.03601v42018
  59. Amortized Set Prediction for Inverse IFS Reconstruction from Density Maps

    Yutaka Yamaguti

    cs.CVarXiv:2608.24175v12026
  60. PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

    Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty +1

    cs.IRcs.CLcs.CVarXiv:2608.26091v12026