Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,221 to 5,280 of 18,866
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
NextStep Team, Chunrui Han, Guopeng Li +47
cs.CVarXiv:2508.10711v22025Asymmetric Bilateral Motion Estimation for Video Frame Interpolation
Junheum Park, Chul Lee, Chang-Su Kim
cs.CVarXiv:2108.06815v12021FLAVR: Flow-Agnostic Video Representations for Fast Frame Interpolation
Tarun Kalluri, Deepak Pathak, Manmohan Chandraker +1
cs.CVarXiv:2012.08512v32020Multi-Modal Answer Validation for Knowledge-Based VQA
Jialin Wu, Jiasen Lu, Ashish Sabharwal +1
cs.CVcs.CLarXiv:2103.12248v32021On the Effectiveness of Image Rotation for Open Set Domain Adaptation
Silvia Bucci, Mohammad Reza Loghmani, Tatiana Tommasi
cs.CVarXiv:2007.12360v12020WoW: Towards a World omniscient World model Through Embodied Interaction
Xiaowei Chi, Peidong Jia, Chun-Kai Fan +33
cs.ROcs.CVcs.MMarXiv:2509.22642v22025Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods
Mehrdad Fazli, Sina Mansouri, Mohit Marvania +1
cs.CVarXiv:2609.01888v12026KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
Yongliang Wu, Zonghui Li, Xinting Hu +7
cs.CVarXiv:2505.16707v12025Ola: Pushing the Frontiers of Omni-Modal Language Model
Zuyan Liu, Yuhao Dong, Jiahui Wang +4
cs.CVcs.CLcs.MMarXiv:2502.04328v32025Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Zhe Kong, Feng Gao, Yong Zhang +5
cs.CVarXiv:2505.22647v12025Accuracy comparison across face recognition algorithms: Where are we on measuring race bias?
Jacqueline G. Cavazos, P. Jonathon Phillips, Carlos D. Castillo +1
cs.CVcs.LGarXiv:1912.07398v22019Dual Graph Convolutional Network for Semantic Segmentation
Li Zhang, Xiangtai Li, Anurag Arnab +3
cs.CVarXiv:1909.06121v32019LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
Hongyu Li, Jinyu Chen, Ziyu Wei +5
cs.CVarXiv:2501.08282v22025ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Xingyu Fu, Minqian Liu, Zhengyuan Yang +6
cs.CVcs.CLarXiv:2501.05452v12025A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
Wenxuan Wang, Zizhan Ma, Zheng Wang +5
cs.CLcs.AIcs.CVarXiv:2502.11211v22025Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
NVIDIA, :, Hassan Abu Alhaija +38
cs.CVcs.AIcs.LGarXiv:2503.14492v22025Selective Structured State-Spaces for Long-Form Video Understanding
Jue Wang, Wentao Zhu, Pichao Wang +4
cs.CVarXiv:2303.14526v12023InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
Shuai Yuan, Yantai Yang, Xiaotian Yang +4
cs.CVarXiv:2601.02281v12026Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
cs.AIcs.CLcs.CVarXiv:2506.09344v12025FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning
Nam Hyeon-Woo, Moon Ye-Bin, Tae-Hyun Oh
cs.LGcs.CVarXiv:2108.06098v32021StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
Jinhui Ye, Ning Gao, Senqiao Yang +7
cs.ROcs.AIcs.CVarXiv:2604.11757v12026LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Komal Kumar, Tajamul Ashraf, Omkar Thawakar +7
cs.CLcs.CVarXiv:2502.21321v22025A CNN-based methodology for breast cancer diagnosis using thermal images
Juan Zuluaga-Gomez, Zeina Al Masry, Khaled Benaggoune +2
cs.CVeess.IVarXiv:1910.13757v12019Real-Time Scene-Adaptive Tone Mapping for High-Dynamic Range Object Detection
Gongzhe Li, Linwei Qiu, Peibei Cao +3
cs.CVarXiv:2608.30400v12026SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations
Yiming Luo, Rongqiang Zhao, Jie Liu
cs.LGcs.CVarXiv:2609.01051v12026Semi-Dense 3D Reconstruction with a Stereo Event Camera
Yi Zhou, Guillermo Gallego, Henri Rebecq +3
cs.CVcs.ROarXiv:1807.07429v12018StreamForest: Efficient Online Video Understanding with Persistent Event Memory
Xiangyu Zeng, Kefan Qiu, Qingyu Zhang +9
cs.CVarXiv:2509.24871v12025Automated Movie Generation via Multi-Agent CoT Planning
Weijia Wu, Zeyu Zhu, Mike Zheng Shou
cs.CVarXiv:2503.07314v12025DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
Zhenjie Yang, Yilin Chai, Xiaosong Jia +5
cs.CVcs.AIcs.ROarXiv:2505.16278v22025Inverse Rendering for Modeling with Line Primitives
Kenji Tojo, Ariel Shamir, Nobuyuki Umetani +1
cs.GRcs.CVarXiv:2609.00625v12026MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation
Zhijian Qiao, Xinjiang Wang, Jiajie Chen +5
cs.CVcs.ROarXiv:2609.01252v12026Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Yunxin Li, Zhenyu Liu, Zitao Li +19
cs.CVcs.CLarXiv:2505.04921v22025NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors
Congyue Deng, Chiyu "Max'' Jiang, Charles R. Qi +4
cs.CVarXiv:2212.03267v12022Wavelet Diffusion Models are fast and scalable Image Generators
Hao Phung, Quan Dao, Anh Tran
cs.CVeess.IVarXiv:2211.16152v22022VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
Yuxuan Bian, Zhaoyang Zhang, Xuan Ju +4
cs.CVcs.AIcs.MMarXiv:2503.05639v32025Streaming Long Video Understanding with Large Language Models
Rui Qian, Xiaoyi Dong, Pan Zhang +4
cs.CVarXiv:2405.16009v12024Improving Object Localization with Fitness NMS and Bounded IoU Loss
Lachlan Tychsen-Smith, Lars Petersson
cs.CVarXiv:1711.00164v32017Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
Zhewei Yao, Amir Gholami, Qi Lei +2
cs.CVcs.LGstat.MLarXiv:1802.08241v42018Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
Zhihe Yang, Xufang Luo, Dongqi Han +2
cs.CVarXiv:2501.09695v22025MASTER: Multi-Aspect Non-local Network for Scene Text Recognition
Ning Lu, Wenwen Yu, Xianbiao Qi +4
cs.CVarXiv:1910.02562v320193DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code
Yipeng Gao, Lei Shu, Genzhi Ye +5
cs.CVcs.AIcs.GRarXiv:2606.01057v12026Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
Nuria Alina Chandra, Hannah Lee, Ryan Murtfeldt +10
cs.CVcs.AIcs.CYarXiv:2503.02857v52025LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent
Jianing Yang, Xuweiyi Chen, Shengyi Qian +4
cs.CVcs.AIcs.CLarXiv:2309.12311v12023PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Soroush Nasiriany, Fei Xia, Wenhao Yu +20
cs.ROcs.CLcs.CVarXiv:2402.07872v12024Pruning and Quantization for Deep Neural Network Acceleration: A Survey
Tailin Liang, John Glossner, Lei Wang +2
cs.CVcs.AIarXiv:2101.09671v32021Transparency of Deep Neural Networks for Medical Image Analysis: A Review of Interpretability Methods
Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee +1
eess.IVcs.AIcs.CVarXiv:2111.02398v12021FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery
Xian Sun, Peijin Wang, Zhiyuan Yan +11
cs.CVarXiv:2103.05569v22021VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
Hiroto Nakata, Yawen Zou, Shunsuke Sakai +5
cs.CVarXiv:2603.13964v12026Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
Yumi Lee, Harim Oh, Hyoryung Kim +52
cs.CVcs.AIarXiv:2609.00866v12026Video models are zero-shot learners and reasoners
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6
cs.LGcs.AIcs.CVarXiv:2509.20328v22025World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
cs.CVcs.AIcs.LGarXiv:2511.00062v22025VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Junxiang Xu, Ruisi Wang, Fanyi Pu +49
cs.CVcs.AIcs.LGarXiv:2608.26105v12026Stitched Value Model for Diffusion Alignment
Hyojun Go, Hyungjin Chung, Prune Truong +8
cs.CVcs.AIcs.LGarXiv:2605.19804v12026Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Shuo Liang, Yixing Ma, Pengfei Zhou +32
cs.CVcs.AIarXiv:2608.14391v12026ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Fan Jiang, Zhaoxu Sun, Mengchao Wang +38
cs.CVcs.AIcs.LGarXiv:2607.19191v12026SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Yuyang Zhao, Yicheng Pan, Qiyuan He +6
cs.CVcs.AIarXiv:2605.30409v12026PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation
Yu Tian, Xintong Jiang, Jan Franklin Adamowski +2
cs.CVarXiv:2609.02860v12026Deep Visual Domain Adaptation: A Survey
Mei Wang, Weihong Deng
cs.CVarXiv:1802.03601v42018Amortized Set Prediction for Inverse IFS Reconstruction from Density Maps
Yutaka Yamaguti
cs.CVarXiv:2608.24175v12026PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty +1
cs.IRcs.CLcs.CVarXiv:2608.26091v12026