Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,441 to 4,500 of 18,830

  1. Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

    Andy Zhou, Bo Li, Haohan Wang

    cs.LGcs.AIcs.CLarXiv:2401.17263v52024
  2. AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

    Xiao Wang, Qingyi Si, Jianlong Wu +3

    cs.CVcs.CLcs.MMarXiv:2503.12559v22025
  3. Residual Connections Encourage Iterative Inference

    Stanisław Jastrzębski, Devansh Arpit, Nicolas Ballas +3

    cs.CVarXiv:1710.04773v22017
  4. Automated Design of Deep Learning Methods for Biomedical Image Segmentation

    Fabian Isensee, Paul F. Jäger, Simon A. A. Kohl +2

    cs.CVarXiv:1904.08128v22019
  5. MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation

    Zhenyu Wu, Yuheng Zhou, Xiuwei Xu +2

    cs.ROcs.CVarXiv:2503.13446v12025
  6. Training-free and Adaptive Sparse Attention for Efficient Long Video Generation

    Yifei Xia, Suhan Ling, Fangcheng Fu +4

    cs.CVarXiv:2502.21079v12025
  7. CineForge: Self-Improving Agents for Long-Horizon Video Generation

    Junxiang Liu, Lin Wang, Haiyu Shi +10

    cs.CVcs.AIarXiv:2608.29621v12026
  8. DeepEMD: Differentiable Earth Mover's Distance for Few-Shot Learning

    Chi Zhang, Yujun Cai, Guosheng Lin +1

    cs.CVcs.LGeess.IVarXiv:2003.06777v52020
  9. State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

    Mingxu Chai, Chenyu Liu, Ziyu Shen +7

    cs.CVarXiv:2608.28698v12026
  10. Progressive LiDAR Adaptation for Road Detection

    Zhe Chen, Jing Zhang, Dacheng Tao

    cs.CVarXiv:1904.01206v12019
  11. A Comprehensive Review of Techniques, Algorithms, Advancements, Challenges, and Clinical Applications of Multi-modal Medical Image Fusion for Improved Diagnosis

    Muhammad Zubair, Muzammil Hussai, Mousa Ahmad Al-Bashrawi +2

    eess.IVcs.CVarXiv:2505.14715v12025
  12. COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval

    Xirong Li, Chaoxi Xu, Xiaoxu Wang +4

    cs.CLcs.CVarXiv:1805.08661v22018
  13. Hydra: A Navigation World Action Model with Discrete Latent Planning and Continuous Flow-Matching Execution

    Mohammad Nazeri, Alexandyr Card, Samira Huber +6

    cs.ROcs.CVarXiv:2608.28995v12026
  14. Protecting Celebrities from DeepFake with Identity Consistency Transformer

    Xiaoyi Dong, Jianmin Bao, Dongdong Chen +6

    cs.CVarXiv:2203.01318v32022
  15. EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation

    Sihang Jia, Shuliang Liu, Songbo Yang +1

    cs.AIcs.CVarXiv:2608.29092v12026
  16. TLA: Tactile-Language-Action Model for Contact-Rich Manipulation

    Peng Hao, Chaofan Zhang, Dingzhe Li +4

    cs.ROcs.CVarXiv:2503.08548v12025
  17. EML-NET:An Expandable Multi-Layer NETwork for Saliency Prediction

    Sen Jia, Neil D. B. Bruce

    cs.CVarXiv:1805.01047v22018
  18. Fast Training of Diffusion Models with Masked Transformers

    Hongkai Zheng, Weili Nie, Arash Vahdat +1

    cs.CVcs.AIcs.LGarXiv:2306.09305v22023
  19. Learning Domain-Invariant Subspace using Domain Features and Independence Maximization

    Ke Yan, Lu Kou, David Zhang

    cs.CVcs.AIcs.LGarXiv:1603.04535v22016
  20. Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

    Xiang Fang, Wanlong Fang, Changshuo Wang

    cs.ROcs.CVarXiv:2606.01565v12026
  21. UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens

    Ruichuan An, Sihan Yang, Renrui Zhang +10

    cs.CVarXiv:2505.14671v32025
  22. GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras

    Ye Yuan, Umar Iqbal, Pavlo Molchanov +2

    cs.CVcs.AIcs.GRarXiv:2112.01524v22021
  23. Unsupervised Prompt Learning for Vision-Language Models

    Tony Huang, Jack Chu, Fangyun Wei

    cs.CVarXiv:2204.03649v22022
  24. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    Mang Ye, Xuankun Rong, Wenke Huang +3

    cs.CRcs.CVarXiv:2502.14881v12025
  25. Neural Graph Matching Network: Learning Lawler's Quadratic Assignment Problem with Extension to Hypergraph and Multiple-graph Matching

    Runzhong Wang, Junchi Yan, Xiaokang Yang

    cs.LGcs.CVstat.MLarXiv:1911.11308v32019
  26. Deep Multimodal Subspace Clustering Networks

    Mahdi Abavisani, Vishal M. Patel

    cs.LGcs.AIcs.CVarXiv:1804.06498v32018
  27. Reformulating HOI Detection as Adaptive Set Prediction

    Mingfei Chen, Yue Liao, Si Liu +3

    cs.CVarXiv:2103.05983v22021
  28. PACO: Parts and Attributes of Common Objects

    Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic +11

    cs.CVarXiv:2301.01795v12023
  29. RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets

    Isabella Liu, Zhan Xu, Wang Yifan +5

    cs.CVarXiv:2502.09615v22025
  30. AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding

    Jiahong Wu, He Zheng, Bo Zhao +9

    cs.CVarXiv:1711.06475v12017
  31. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

    Can Cui, Pengxiang Ding, Wenxuan Song +10

    cs.ROcs.CVarXiv:2505.03912v12025
  32. Panacea: Panoramic and Controllable Video Generation for Autonomous Driving

    Yuqing Wen, Yucheng Zhao, Yingfei Liu +7

    cs.CVarXiv:2311.16813v12023
  33. Unsupervised Tracklet Person Re-Identification

    Minxian Li, Xiatian Zhu, Shaogang Gong

    cs.CVarXiv:1903.00535v12019
  34. Multi-source weak supervision for saliency detection

    Yu Zeng, Yunzhi Zhuge, Huchuan Lu +3

    cs.CVarXiv:1904.00566v12019
  35. Token Contrast for Weakly-Supervised Semantic Segmentation

    Lixiang Ru, Heliang Zheng, Yibing Zhan +1

    cs.CVarXiv:2303.01267v12023
  36. HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

    Jiaxing Zhao, Qize Yang, Yixing Peng +8

    cs.CVarXiv:2501.15111v12025
  37. Action Quality Assessment Across Multiple Actions

    Paritosh Parmar, Brendan Tran Morris

    cs.CVarXiv:1812.06367v22018
  38. Invisible Image Watermarks Are Provably Removable Using Generative AI

    Xuandong Zhao, Kexun Zhang, Zihao Su +6

    cs.CRcs.AIcs.CVarXiv:2306.01953v32023
  39. Recurrent Pixel Embedding for Instance Grouping

    Shu Kong, Charless Fowlkes

    cs.CVcs.LGcs.MMarXiv:1712.08273v12017
  40. Adaptive Frequency Enhancement Network for Remote Sensing Image Semantic Segmentation

    Feng Gao, Miao Fu, Jingchao Cao +2

    eess.IVcs.CVarXiv:2504.02647v12025
  41. Lizard: A Large-Scale Dataset for Colonic Nuclear Instance Segmentation and Classification

    Simon Graham, Mostafa Jahanifar, Ayesha Azam +14

    cs.CVcs.LGarXiv:2108.11195v22021
  42. Leveraging Shape Completion for 3D Siamese Tracking

    Silvio Giancola, Jesus Zarzar, Bernard Ghanem

    cs.CVarXiv:1903.01784v22019
  43. Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

    Yifan Du, Zikang Liu, Yifan Li +7

    cs.CVcs.AIarXiv:2501.01904v22025
  44. Joint Face Alignment and 3D Face Reconstruction with Application to Face Recognition

    Feng Liu, Qijun Zhao, Xiaoming Liu +1

    cs.CVarXiv:1708.02734v22017
  45. Comparison of Various SLAM Systems for Mobile Robot in an Indoor Environment

    Maksim Filipenko, Ilya Afanasyev

    cs.ROcs.CVarXiv:2501.09490v12025
  46. Intuitive physics understanding emerges from self-supervised pretraining on natural videos

    Quentin Garrido, Nicolas Ballas, Mahmoud Assran +5

    cs.CVcs.AIarXiv:2502.11831v12025
  47. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

    Yuhang Zang, Xiaoyi Dong, Pan Zhang +10

    cs.CVcs.CLarXiv:2501.12368v22025
  48. Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond

    Guanyao Wu, Haoyu Liu, Hongming Fu +4

    cs.CVarXiv:2503.01210v22025
  49. Linkage Based Face Clustering via Graph Convolution Network

    Zhongdao Wang, Liang Zheng, Yali Li +1

    cs.CVarXiv:1903.11306v32019
  50. Opportunities and Challenges of Deep Learning Methods for Electrocardiogram Data: A Systematic Review

    Shenda Hong, Yuxi Zhou, Junyuan Shang +2

    eess.SPcs.CVcs.LGarXiv:2001.01550v32019
  51. A Survey on Deep Learning for Skin Lesion Segmentation

    Zahra Mirikharaji, Kumar Abhishek, Alceu Bissoto +5

    eess.IVcs.CVcs.LGarXiv:2206.00356v32022
  52. RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization

    Songming Liu, Bangguo Li, Kai Ma +5

    cs.ROcs.AIcs.CVarXiv:2602.03310v12026
  53. CE-FPN: Enhancing Channel Information for Object Detection

    Yihao Luo, Xiang Cao, Juntao Zhang +5

    cs.CVarXiv:2103.10643v12021
  54. Face Recognition Methods & Applications

    Divyarajsinh N. Parmar, Brijesh B. Mehta

    cs.CVarXiv:1403.0485v12014
  55. Graph2Plan: Learning Floorplan Generation from Layout Graphs

    Ruizhen Hu, Zeyu Huang, Yuhan Tang +3

    cs.CVcs.GRarXiv:2004.13204v12020
  56. SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training

    Jintao Zhang, Jia Wei, Pengle Zhang +6

    cs.LGcs.AIcs.ARarXiv:2505.11594v32025
  57. Intelligent Approaches to interact with Machines using Hand Gesture Recognition in Natural way: A Survey

    Ankit Chaudhary, J. L. Raheja, Karen Das +1

    cs.HCcs.CVarXiv:1303.2292v12013
  58. VLP: Vision Language Planning for Autonomous Driving

    Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti +4

    cs.CVarXiv:2401.05577v42024
  59. A Performance Analysis of You Only Look Once Models for Deployment on Constrained Computational Edge Devices in Drone Applications

    Lucas Rey, Ana M. Bernardos, Andrzej D. Dobrzycki +4

    cs.DCcs.AIcs.CVarXiv:2502.15737v12025
  60. DRPCA-Net: Make Robust PCA Great Again for Infrared Small Target Detection

    Zihao Xiong, Fei Zhou, Fengyi Wu +5

    cs.CVarXiv:2507.09541v12025