Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,141 to 16,200 of 18,867

  1. MediX-R1: Open Ended Medical Reinforcement Learning

    Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed +5

    cs.CVarXiv:2602.23363v12026
  2. ImageNet-21K Pretraining for the Masses

    Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy +1

    cs.CVcs.LGarXiv:2104.10972v42021
  3. Multi-attentional Deepfake Detection

    Hanqing Zhao, Wenbo Zhou, Dongdong Chen +3

    cs.CVarXiv:2103.02406v32021
  4. MoCha:End-to-End Video Character Replacement without Structural Guidance

    Zhengbo Xu, Jie Ma, Ziheng Wang +3

    cs.CVarXiv:2601.08587v22026
  5. SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

    Zirui Wang, Jiahui Yu, Adams Wei Yu +3

    cs.CVcs.CLcs.LGarXiv:2108.10904v32021
  6. Going Deeper in Facial Expression Recognition using Deep Neural Networks

    Ali Mollahosseini, David Chan, Mohammad H. Mahoor

    cs.NEcs.CVarXiv:1511.04110v12015
  7. Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction

    Ziyi Yang, Xinyu Gao, Wen Zhou +3

    cs.CVarXiv:2309.13101v22023
  8. Action100M: A Large-scale Video Action Dataset

    Delong Chen, Tejaswi Kasarla, Yejin Bang +6

    cs.CVarXiv:2601.10592v12026
  9. Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures

    Hengyuan Hu, Rui Peng, Yu-Wing Tai +1

    cs.NEcs.CVcs.LGarXiv:1607.03250v12016
  10. Image-Image Domain Adaptation with Preserved Self-Similarity and Domain-Dissimilarity for Person Re-identification

    Weijian Deng, Liang Zheng, Qixiang Ye +3

    cs.CVarXiv:1711.07027v32017
  11. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

    Lijiang Li, Zuwei Long, Yunhang Shen +6

    cs.CVarXiv:2603.06577v22026
  12. UIU-Net: U-Net in U-Net for Infrared Small Object Detection

    Xin Wu, Danfeng Hong, Jocelyn Chanussot

    cs.CVarXiv:2212.00968v12022
  13. SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines

    Yinda Xu, Zeyu Wang, Zuoxin Li +2

    cs.CVarXiv:1911.06188v42019
  14. The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes

    Douwe Kiela, Hamed Firooz, Aravind Mohan +4

    cs.AIcs.CLcs.CVarXiv:2005.04790v32020
  15. RMA: Rapid Motor Adaptation for Legged Robots

    Ashish Kumar, Zipeng Fu, Deepak Pathak +1

    cs.LGcs.AIcs.CVarXiv:2107.04034v12021
  16. Large Batch Training of Convolutional Networks

    Yang You, Igor Gitman, Boris Ginsburg

    cs.CVarXiv:1708.03888v32017
  17. Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation

    George Papandreou, Liang-Chieh Chen, Kevin Murphy +1

    cs.CVarXiv:1502.02734v32015
  18. A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark +2

    cs.CVcs.CLarXiv:2206.01718v12022
  19. The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

    Xiaoze Liu, Ruowang Zhang, Weichen Yu +7

    cs.CLcs.CVcs.LGarXiv:2602.15382v22026
  20. MWM: Mobile World Models for Action-Conditioned Consistent Prediction

    Han Yan, Zishang Xiang, Zeyu Zhang +1

    cs.CVcs.ROarXiv:2603.07799v12026
  21. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising

    Kai Zou, Dian Zheng, Hongbo Liu +3

    cs.CVarXiv:2603.08703v12026
  22. LIVE: Long-horizon Interactive Video World Modeling

    Junchao Huang, Ziyang Ye, Xinting Hu +5

    cs.CVarXiv:2602.03747v12026
  23. Temporal Segment Networks for Action Recognition in Videos

    Limin Wang, Yuanjun Xiong, Zhe Wang +4

    cs.CVarXiv:1705.02953v12017
  24. NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions

    Junbin Xiao, Xindi Shang, Angela Yao +1

    cs.CVcs.AIarXiv:2105.08276v22021
  25. pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis

    Eric R. Chan, Marco Monteiro, Petr Kellnhofer +2

    cs.CVcs.GRarXiv:2012.00926v22020
  26. Medical Image Segmentation Review: The success of U-Net

    Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland +7

    eess.IVcs.CVarXiv:2211.14830v12022
  27. LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model

    Quankai Gao, Jiawei Yang, Qiangeng Xu +2

    cs.CVarXiv:2603.27449v12026
  28. Mobile-GS: Real-time Gaussian Splatting for Mobile Devices

    Xiaobiao Du, Yida Wang, Kun Zhan +1

    cs.CVarXiv:2603.11531v12026
  29. ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models

    Jooyoung Choi, Sungwon Kim, Yonghyun Jeong +2

    cs.CVarXiv:2108.02938v22021
  30. Understanding data augmentation for classification: when to warp?

    Sebastien C. Wong, Adam Gatt, Victor Stamatescu +1

    cs.CVarXiv:1609.08764v22016
  31. DeepID3: Face Recognition with Very Deep Neural Networks

    Yi Sun, Ding Liang, Xiaogang Wang +1

    cs.CVarXiv:1502.00873v12015
  32. Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

    Gen Li, Nan Duan, Yuejian Fang +3

    cs.CVarXiv:1908.06066v32019
  33. Making Convolutional Networks Shift-Invariant Again

    Richard Zhang

    cs.CVcs.LGarXiv:1904.11486v22019
  34. PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference

    Xiaofeng Mao, Shaohao Rui, Kaining Ying +4

    cs.CVcs.AIarXiv:2603.25730v12026
  35. Sparsity Invariant CNNs

    Jonas Uhrig, Nick Schneider, Lukas Schneider +3

    cs.CVarXiv:1708.06500v22017
  36. Semi-Supervised Semantic Segmentation with Cross-Consistency Training

    Yassine Ouali, Céline Hudelot, Myriam Tami

    cs.CVarXiv:2003.09005v32020
  37. MAttNet: Modular Attention Network for Referring Expression Comprehension

    Licheng Yu, Zhe Lin, Xiaohui Shen +4

    cs.CVcs.AIcs.CLarXiv:1801.08186v32018
  38. Zero-Shot Learning -- The Good, the Bad and the Ugly

    Yongqin Xian, Bernt Schiele, Zeynep Akata

    cs.CVarXiv:1703.04394v22017
  39. Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching

    Xiaodong Gu, Zhiwen Fan, Zuozhuo Dai +3

    cs.CVarXiv:1912.06378v32019
  40. Deep Fragment Embeddings for Bidirectional Image Sentence Mapping

    Andrej Karpathy, Armand Joulin, Li Fei-Fei

    cs.CVcs.CLcs.LGarXiv:1406.5679v12014
  41. TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

    Bin Yu, Shijie Lian, Xiaopeng Lin +8

    cs.ROcs.CVarXiv:2601.14133v22026
  42. MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

    Yanghao Li, Chao-Yuan Wu, Haoqi Fan +4

    cs.CVarXiv:2112.01526v22021
  43. KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D

    Yiyi Liao, Jun Xie, Andreas Geiger

    cs.CVarXiv:2109.13410v22021
  44. BoT-SORT: Robust Associations Multi-Pedestrian Tracking

    Nir Aharon, Roy Orfaig, Ben-Zion Bobrovsky

    cs.CVarXiv:2206.14651v22022
  45. SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

    Boyuan Chen, Zhuo Xu, Sean Kirmani +6

    cs.CVcs.CLcs.LGarXiv:2401.12168v12024
  46. Repurposing Geometric Foundation Models for Multi-view Diffusion

    Wooseok Jang, Seonghu Jeon, Jisang Han +5

    cs.CVarXiv:2603.22275v12026
  47. CLIPort: What and Where Pathways for Robotic Manipulation

    Mohit Shridhar, Lucas Manuelli, Dieter Fox

    cs.ROcs.CLcs.CVarXiv:2109.12098v12021
  48. MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation

    Yejin Kim, Wilbert Pumacay, Omar Rayyan +23

    cs.ROcs.AIcs.CVarXiv:2602.11337v22026
  49. One-Shot Video Object Segmentation

    Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset +3

    cs.CVarXiv:1611.05198v42016
  50. Translating Videos to Natural Language Using Deep Recurrent Neural Networks

    Subhashini Venugopalan, Huijuan Xu, Jeff Donahue +3

    cs.CVcs.CLarXiv:1412.4729v32014
  51. Designing a Practical Degradation Model for Deep Blind Image Super-Resolution

    Kai Zhang, Jingyun Liang, Luc Van Gool +1

    eess.IVcs.CVarXiv:2103.14006v22021
  52. SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

    Bohao Li, Rui Wang, Guangzhi Wang +3

    cs.CLcs.CVarXiv:2307.16125v22023
  53. Temporal Action Detection with Structured Segment Networks

    Yue Zhao, Yuanjun Xiong, Limin Wang +3

    cs.CVarXiv:1704.06228v22017
  54. "Zero-Shot" Super-Resolution using Deep Internal Learning

    Assaf Shocher, Nadav Cohen, Michal Irani

    cs.CVcs.LGcs.NEarXiv:1712.06087v12017
  55. Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose

    Georgios Pavlakos, Xiaowei Zhou, Konstantinos G. Derpanis +1

    cs.CVarXiv:1611.07828v22016
  56. Leveraging Frequency Analysis for Deep Fake Image Recognition

    Joel Frank, Thorsten Eisenhofer, Lea Schönherr +3

    cs.CVeess.IVarXiv:2003.08685v32020
  57. Land-Cover Classification with High-Resolution Remote Sensing Images Using Transferable Deep Models

    Xin-Yi Tong, Gui-Song Xia, Qikai Lu +4

    cs.CVarXiv:1807.05713v32018
  58. Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

    Meiqi Wu, Zhixin Cai, Fufangchen Zhao +13

    cs.CVarXiv:2603.22212v12026
  59. GLIGEN: Open-Set Grounded Text-to-Image Generation

    Yuheng Li, Haotian Liu, Qingyang Wu +5

    cs.CVcs.AIcs.CLarXiv:2301.07093v22023
  60. Invariant Information Clustering for Unsupervised Image Classification and Segmentation

    Xu Ji, João F. Henriques, Andrea Vedaldi

    cs.CVcs.LGarXiv:1807.06653v42018