Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,401 to 5,460 of 18,855

  1. Layered Neural Atlases for Consistent Video Editing

    Yoni Kasten, Dolev Ofri, Oliver Wang +1

    cs.CVcs.GRarXiv:2109.11418v12021
  2. SurgSkill-Bench: A Benchmark for Multimodal Surgical Skill Assessment

    Chaohui Dang, Zheheng Jiang, James Glasbey +3

    cs.CVarXiv:2608.30872v12026
  3. ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

    Adrien Mialland, Marc Plantevit, Julien Gallois +1

    cs.IRcs.AIcs.CLarXiv:2609.02486v12026
  4. Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation

    Chenjie Cao, Jingkai Zhou, Shikai Li +5

    cs.CVarXiv:2504.14899v22025
  5. Anomaly Detection in Video Using Predictive Convolutional Long Short-Term Memory Networks

    Jefferson Ryan Medel, Andreas Savakis

    cs.CVarXiv:1612.00390v22016
  6. UnionDet: Union-Level Detector Towards Real-Time Human-Object Interaction Detection

    Bumsoo Kim, Taeho Choi, Jaewoo Kang +1

    cs.CVarXiv:2312.12664v12023
  7. PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices

    Guanghua Yu, Qinyao Chang, Wenyu Lv +12

    cs.CVarXiv:2111.00902v12021
  8. Diffuse and Disperse: Image Generation with Representation Regularization

    Runqian Wang, Kaiming He

    cs.CVcs.AIcs.LGarXiv:2506.09027v22025
  9. Defeating Image Obfuscation with Deep Learning

    Richard McPherson, Reza Shokri, Vitaly Shmatikov

    cs.CRcs.CVarXiv:1609.00408v22016
  10. Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

    Ke Fan, Shunlin Lu, Minyue Dai +6

    cs.CVarXiv:2507.07095v12025
  11. An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

    Weike Zhao, Chaoyi Wu, Yanjie Fan +10

    cs.CLcs.AIcs.CVarXiv:2506.20430v32025
  12. Learning Joint 2D-3D Representations for Depth Completion

    Yun Chen, Bin Yang, Ming Liang +1

    cs.CVarXiv:2012.12402v12020
  13. How to build a consistency model: Learning flow maps via self-distillation

    Nicholas M. Boffi, Michael S. Albergo, Eric Vanden-Eijnden

    cs.LGcs.CVarXiv:2505.18825v22025
  14. Global Sparse Momentum SGD for Pruning Very Deep Neural Networks

    Xiaohan Ding, Guiguang Ding, Xiangxin Zhou +3

    cs.LGcs.CVstat.MLarXiv:1909.12778v32019
  15. SkyReels-A2: Compose Anything in Video Diffusion Transformers

    Zhengcong Fei, Debang Li, Di Qiu +8

    cs.CVarXiv:2504.02436v12025
  16. Uncertainty-Aware Trajectory Forecasting from Imperfect Tracking

    Stephane Da Silva Martins, Victor Petrovic, Emanuel Aldea +1

    cs.CVarXiv:2608.30899v12026
  17. Information-Flow Matting

    Yağız Aksoy, Tunç Ozan Aydın, Marc Pollefeys

    cs.CVarXiv:1707.05055v22017
  18. M-LLM Based Video Frame Selection for Efficient Video Understanding

    Kai Hu, Feng Gao, Xiaohan Nie +8

    cs.CVcs.AIarXiv:2502.19680v22025
  19. PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation

    Haibo Jin, Haoxuan Che, Yi Lin +1

    cs.CVcs.CLarXiv:2308.12604v22023
  20. DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

    Ximeng Sun, Ping Hu, Kate Saenko

    cs.CVarXiv:2206.09541v12022
  21. MASA-SR: Matching Acceleration and Spatial Adaptation for Reference-Based Image Super-Resolution

    Liying Lu, Wenbo Li, Xin Tao +2

    cs.CVarXiv:2106.02299v12021
  22. Neural 3D Morphable Models: Spiral Convolutional Networks for 3D Shape Representation Learning and Generation

    Giorgos Bouritsas, Sergiy Bokhnyak, Stylianos Ploumpis +2

    cs.CVcs.AIcs.GRarXiv:1905.02876v32019
  23. dhSegment: A generic deep-learning approach for document segmentation

    Sofia Ares Oliveira, Benoit Seguin, Frederic Kaplan

    cs.CVarXiv:1804.10371v22018
  24. SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation

    Hang Yin, Xiuwei Xu, Zhenyu Wu +2

    cs.CVcs.ROarXiv:2410.08189v12024
  25. Re-evaluating Automatic Metrics for Image Captioning

    Mert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis +1

    cs.CLcs.CVarXiv:1612.07600v12016
  26. Group-aware Label Transfer for Domain Adaptive Person Re-identification

    Kecheng Zheng, Wu Liu, Lingxiao He +3

    cs.CVarXiv:2103.12366v12021
  27. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

    Raul Ortega, José Manuel Gómez-Pérez

    cs.CVcs.AIcs.CLarXiv:2609.00948v12026
  28. RVT-2: Learning Precise Manipulation from Few Demonstrations

    Ankit Goyal, Valts Blukis, Jie Xu +3

    cs.ROcs.AIcs.CVarXiv:2406.08545v12024
  29. StreamingVLM: Real-Time Understanding for Infinite Video Streams

    Ruyi Xu, Guangxuan Xiao, Yukang Chen +3

    cs.CVcs.AIcs.CLarXiv:2510.09608v22025
  30. XingGAN for Person Image Generation

    Hao Tang, Song Bai, Li Zhang +2

    cs.CVcs.LGeess.IVarXiv:2007.09278v12020
  31. CHIP: CHannel Independence-based Pruning for Compact Neural Networks

    Yang Sui, Miao Yin, Yi Xie +3

    cs.CVcs.AIcs.LGarXiv:2110.13981v32021
  32. Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

    Seil Kang, Jinyeong Kim, Junhyeok Kim +1

    cs.CVcs.AIarXiv:2503.06287v12025
  33. Latent Replay for Real-Time Continual Learning

    Lorenzo Pellegrini, Gabriele Graffieti, Vincenzo Lomonaco +1

    cs.LGcs.CVstat.MLarXiv:1912.01100v22019
  34. You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change

    Kaizhen Tan

    cs.CVarXiv:2609.00649v12026
  35. UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

    Sule Bai, Mingxing Li, Yong Liu +5

    cs.CVarXiv:2505.14231v12025
  36. SOON: Scenario Oriented Object Navigation with Graph-based Exploration

    Fengda Zhu, Xiwen Liang, Yi Zhu +2

    cs.CVarXiv:2103.17138v22021
  37. VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

    Jing Liu, Sihan Chen, Xingjian He +4

    cs.LGcs.CLcs.CVarXiv:2304.08345v22023
  38. Multi3DRefer: Grounding Text Description to Multiple 3D Objects

    Yiming Zhang, ZeMing Gong, Angel X. Chang

    cs.CVarXiv:2309.05251v12023
  39. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

    Dongzhi Jiang, Renrui Zhang, Ziyu Guo +11

    cs.CVcs.AIcs.CLarXiv:2502.09621v12025
  40. Unsupervised Learning of Long-Term Motion Dynamics for Videos

    Zelun Luo, Boya Peng, De-An Huang +2

    cs.CVarXiv:1701.01821v32017
  41. FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality Assessment

    Jinglin Xu, Yongming Rao, Xumin Yu +3

    cs.CVarXiv:2204.03646v12022
  42. MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm

    Zhang Li, Yuliang Liu, Qiang Liu +8

    cs.CVarXiv:2506.05218v22025
  43. VSA: Faster Video Diffusion with Trainable Sparse Attention

    Peiyuan Zhang, Yongqi Chen, Haofeng Huang +5

    cs.CVarXiv:2505.13389v52025
  44. Patch Diffusion: Faster and More Data-Efficient Training of Diffusion Models

    Zhendong Wang, Yifan Jiang, Huangjie Zheng +5

    cs.CVcs.LGarXiv:2304.12526v22023
  45. Reduced Reference Perceptual Quality Model and Application to Rate Control for 3D Point Cloud Compression

    Qi Liu, Hui Yuan, Raouf Hamzaoui +3

    eess.IVcs.CVarXiv:2011.12688v12020
  46. GameFactory: Creating New Games with Generative Interactive Videos

    Jiwen Yu, Yiran Qin, Xintao Wang +3

    cs.CVarXiv:2501.08325v42025
  47. Deep Audio-Visual Learning: A Survey

    Hao Zhu, Mandi Luo, Rui Wang +2

    cs.CVarXiv:2001.04758v12020
  48. RealOOB: A Definition-Consistent Real-World Oriented Occlusion Boundary Benchmark

    Lintao Xu, Yinghao Wang, Chenchu Rong +2

    cs.CVarXiv:2608.30820v12026
  49. UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

    Duomin Wang, Wei Zuo, Aojie Li +7

    cs.CVarXiv:2509.06155v12025
  50. Scaling RL to Long Videos

    Yukang Chen, Wei Huang, Baifeng Shi +11

    cs.CVcs.AIcs.CLarXiv:2507.07966v42025
  51. Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

    Qingchan Zhu, Weihang You, Hanqi Jiang +3

    cs.CVcs.CLcs.LGarXiv:2609.03158v12026
  52. Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

    Chuer Chen, Zichen Wang, Yi He +2

    cs.CVcs.AIarXiv:2609.02502v12026
  53. Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

    Ziyu Zhu, Xilin Wang, Yixuan Li +9

    cs.CVarXiv:2507.04047v22025
  54. Countering Malicious DeepFakes: Survey, Battleground, and Horizon

    Felix Juefei-Xu, Run Wang, Yihao Huang +3

    cs.CVarXiv:2103.00218v32021
  55. Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition

    Jiaqi Li, Junshu Tang, Zhiyong Xu +6

    cs.CVarXiv:2506.17201v12025
  56. GR-3 Technical Report

    Chilam Cheang, Sijin Chen, Zhongren Cui +18

    cs.ROcs.AIcs.CVarXiv:2507.15493v22025
  57. CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning

    Baraa Bilbeisi, Mengchen Fan, Baocheng Geng +1

    cs.LGcs.CVarXiv:2609.00446v12026
  58. Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

    Xiangyu Zhao, Peiyuan Zhang, Kexian Tang +10

    cs.CVarXiv:2504.02826v42025
  59. Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video Detection

    Siyu Li, Jin Yang, Weiheng Liang

    cs.CVcs.MMarXiv:2609.00742v12026
  60. CT-Realistic Lung Nodule Simulation from 3D Conditional Generative Adversarial Networks for Robust Lung Segmentation

    Dakai Jin, Ziyue Xu, Youbao Tang +2

    cs.CVarXiv:1806.04051v12018