Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,221 to 14,280 of 18,811

  1. Context Encoders: Feature Learning by Inpainting

    Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue +2

    cs.CVcs.AIcs.GRarXiv:1604.07379v22016
  2. Conditional Random Fields as Recurrent Neural Networks

    Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes +5

    cs.CVarXiv:1502.03240v32015
  3. CIDEr: Consensus-based Image Description Evaluation

    Ramakrishna Vedantam, C. Lawrence Zitnick, Devi Parikh

    cs.CVcs.CLcs.IRarXiv:1411.5726v22014
  4. End-to-End Dense Video Captioning with Masked Transformer

    Luowei Zhou, Yingbo Zhou, Jason J. Corso +2

    cs.CVarXiv:1804.00819v12018
  5. End-to-End Comparative Attention Networks for Person Re-identification

    Hao Liu, Jiashi Feng, Meibin Qi +2

    cs.CVarXiv:1606.04404v22016
  6. Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation

    Weizhi Tao, Zengwang Jin, Xiao Wang +1

    cs.CVcs.ROarXiv:2608.24366v12026
  7. Video Salient Object Detection via Fully Convolutional Networks

    Wenguan Wang, Jianbing Shen, Ling Shao

    cs.CVarXiv:1702.00871v32017
  8. StyleFlow: Attribute-conditioned Exploration of StyleGAN-Generated Images using Conditional Continuous Normalizing Flows

    Rameen Abdal, Peihao Zhu, Niloy Mitra +1

    cs.CVcs.GRarXiv:2008.02401v22020
  9. PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection

    Shaoshuai Shi, Li Jiang, Jiajun Deng +5

    cs.CVarXiv:2102.00463v32021
  10. Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision

    Xinchen Yan, Jimei Yang, Ersin Yumer +2

    cs.CVcs.GRcs.LGarXiv:1612.00814v32016
  11. WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation

    Jongheon Jeong, Yang Zou, Taewan Kim +3

    cs.CVcs.AIcs.CLarXiv:2303.14814v12023
  12. PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

    Ziqi Cui, Shangyu Lou

    cs.CVcs.AIcs.IRarXiv:2608.24133v12026
  13. AANet: Adaptive Aggregation Network for Efficient Stereo Matching

    Haofei Xu, Juyong Zhang

    cs.CVarXiv:2004.09548v12020
  14. Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

    Patrick Schramowski, Manuel Brack, Björn Deiseroth +1

    cs.CVcs.AIcs.LGarXiv:2211.05105v42022
  15. LeFlow: Generative Latent Flow Planning for World Models

    Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu +1

    cs.CVarXiv:2608.24855v12026
  16. Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting

    Zeyu Yang, Hongye Yang, Zijie Pan +1

    cs.CVarXiv:2310.10642v32023
  17. MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge

    Linxi Fan, Guanzhi Wang, Yunfan Jiang +7

    cs.LGcs.AIcs.CLarXiv:2206.08853v22022
  18. It Is Not the Journey but the Destination: Endpoint Conditioned Trajectory Prediction

    Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal +4

    cs.CVcs.LGarXiv:2004.02025v32020
  19. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

    Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman +2

    cs.AIcs.CLcs.CVarXiv:2310.04406v32023
  20. Sliced Wasserstein Discrepancy for Unsupervised Domain Adaptation

    Chen-Yu Lee, Tanmay Batra, Mohammad Haris Baig +1

    cs.CVcs.LGstat.MLarXiv:1903.04064v12019
  21. Finding Action Tubes

    Georgia Gkioxari, Jitendra Malik

    cs.CVarXiv:1411.6031v12014
  22. A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts

    Jian Liang, Ran He, Tieniu Tan

    cs.LGcs.AIcs.CVarXiv:2303.15361v22023
  23. Learning to Reason: End-to-End Module Networks for Visual Question Answering

    Ronghang Hu, Jacob Andreas, Marcus Rohrbach +2

    cs.CVarXiv:1704.05526v32017
  24. Spatial-Phase Shallow Learning: Rethinking Face Forgery Detection in Frequency Domain

    Honggu Liu, Xiaodan Li, Wenbo Zhou +5

    cs.CVarXiv:2103.01856v32021
  25. VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

    Haoxin Chen, Menghan Xia, Yingqing He +9

    cs.CVarXiv:2310.19512v12023
  26. Camera Style Adaptation for Person Re-identification

    Zhun Zhong, Liang Zheng, Zhedong Zheng +2

    cs.CVarXiv:1711.10295v22017
  27. Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach

    Xingyi Zhou, Qixing Huang, Xiao Sun +2

    cs.CVarXiv:1704.02447v22017
  28. A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection

    Xiaolong Wang, Abhinav Shrivastava, Abhinav Gupta

    cs.CVarXiv:1704.03414v12017
  29. GPT-4V(ision) is a Generalist Web Agent, if Grounded

    Boyuan Zheng, Boyu Gou, Jihyung Kil +2

    cs.IRcs.AIcs.CLarXiv:2401.01614v22024
  30. OneFormer: One Transformer to Rule Universal Image Segmentation

    Jitesh Jain, Jiachen Li, MangTik Chiu +3

    cs.CVarXiv:2211.06220v22022
  31. SWAD: Domain Generalization by Seeking Flat Minima

    Junbum Cha, Sanghyuk Chun, Kyungjae Lee +4

    cs.LGcs.CVarXiv:2102.08604v42021
  32. Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

    Haoning Wu, Zicheng Zhang, Weixia Zhang +11

    cs.CVcs.CLcs.LGarXiv:2312.17090v12023
  33. Tangent Convolutions for Dense Prediction in 3D

    Maxim Tatarchenko, Jaesik Park, Vladlen Koltun +1

    cs.CVarXiv:1807.02443v12018
  34. Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

    Zhe Jin, Zhimin Lin, Bin Zheng +2

    cs.CVcs.AIarXiv:2608.23011v12026
  35. MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

    Liangtao Shi, Jinxia Xie, Xiantao Hu +1

    cs.CVarXiv:2608.23234v12026
  36. A Twofold Siamese Network for Real-Time Object Tracking

    Anfeng He, Chong Luo, Xinmei Tian +1

    cs.CVarXiv:1802.08817v12018
  37. Region Proposal by Guided Anchoring

    Jiaqi Wang, Kai Chen, Shuo Yang +2

    cs.CVarXiv:1901.03278v22019
  38. Learning 3D Human Dynamics from Video

    Angjoo Kanazawa, Jason Y. Zhang, Panna Felsen +1

    cs.CVarXiv:1812.01601v42018
  39. Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026

    Jinxing Zhou, Suiyi Zhao, Yanghao Zhou +1

    cs.MMcs.CVcs.SDarXiv:2608.22337v12026
  40. Incremental Learning of Object Detectors without Catastrophic Forgetting

    Konstantin Shmelkov, Cordelia Schmid, Karteek Alahari

    cs.CVarXiv:1708.06977v12017
  41. StegaStamp: Invisible Hyperlinks in Physical Photographs

    Matthew Tancik, Ben Mildenhall, Ren Ng

    cs.CVarXiv:1904.05343v22019
  42. Platonic Representation Hypothesis on World Models

    Wenhow Li, Chengwei MA, Hui Xiong +2

    cs.CVarXiv:2608.23720v12026
  43. Learning Data Augmentation Strategies for Object Detection

    Barret Zoph, Ekin D. Cubuk, Golnaz Ghiasi +3

    cs.CVcs.LGarXiv:1906.11172v12019
  44. Deep Global Registration

    Christopher Choy, Wei Dong, Vladlen Koltun

    cs.CVcs.CGcs.LGarXiv:2004.11540v22020
  45. Generalizing Face Forgery Detection with High-frequency Features

    Yuchen Luo, Yong Zhang, Junchi Yan +1

    cs.CVarXiv:2103.12376v12021
  46. MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers

    Huiyu Wang, Yukun Zhu, Hartwig Adam +2

    cs.CVarXiv:2012.00759v32020
  47. Cross-Generation Optimization of YOLOv26, YOLOv11, and YOLOv8 for Fine-Grained Small-Object Detection and Instance Segmentation in Complex Orchards

    Ranjan Sapkota, Manoj Karkee

    cs.CVarXiv:2608.23636v12026
  48. Large-scale Multi-view Subspace Clustering in Linear Time

    Zhao Kang, Wangtao Zhou, Zhitong Zhao +3

    cs.LGcs.CVstat.MLarXiv:1911.09290v12019
  49. PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation

    Yang Zhang, Zixiang Zhou, Philip David +4

    cs.CVarXiv:2003.14032v22020
  50. Accelerating Eulerian Fluid Simulation With Convolutional Networks

    Jonathan Tompson, Kristofer Schlachter, Pablo Sprechmann +1

    cs.CVarXiv:1607.03597v72016
  51. Meta R-CNN : Towards General Solver for Instance-level Few-shot Learning

    Xiaopeng Yan, Ziliang Chen, Anni Xu +3

    cs.CVcs.LGarXiv:1909.13032v22019
  52. Invariance Matters: Exemplar Memory for Domain Adaptive Person Re-identification

    Zhun Zhong, Liang Zheng, Zhiming Luo +2

    cs.CVcs.LGarXiv:1904.01990v12019
  53. Open Set Domain Adaptation by Backpropagation

    Kuniaki Saito, Shohei Yamamoto, Yoshitaka Ushiku +1

    cs.CVarXiv:1804.10427v22018
  54. Low-Light Image Enhancement with Normalizing Flow

    Yufei Wang, Renjie Wan, Wenhan Yang +3

    eess.IVcs.CVarXiv:2109.05923v12021
  55. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language

    Songyang Zhang, Houwen Peng, Jianlong Fu +1

    cs.CVcs.IRcs.MMarXiv:1912.03590v32019
  56. Visual Semantic Reasoning for Image-Text Matching

    Kunpeng Li, Yulun Zhang, Kai Li +2

    cs.CVarXiv:1909.02701v12019
  57. Learning to Compose Dynamic Tree Structures for Visual Contexts

    Kaihua Tang, Hanwang Zhang, Baoyuan Wu +2

    cs.CVarXiv:1812.01880v12018
  58. $A^2$-Nets: Double Attention Networks

    Yunpeng Chen, Yannis Kalantidis, Jianshu Li +2

    cs.CVarXiv:1810.11579v12018
  59. A study of the effect of JPG compression on adversarial images

    Gintare Karolina Dziugaite, Zoubin Ghahramani, Daniel M. Roy

    cs.CVcs.LGarXiv:1608.00853v12016
  60. Asymmetric Tri-training for Unsupervised Domain Adaptation

    Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada

    cs.CVcs.AIarXiv:1702.08400v32017