Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,221 to 14,280 of 18,811
Context Encoders: Feature Learning by Inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue +2
cs.CVcs.AIcs.GRarXiv:1604.07379v22016Conditional Random Fields as Recurrent Neural Networks
Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes +5
cs.CVarXiv:1502.03240v32015CIDEr: Consensus-based Image Description Evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, Devi Parikh
cs.CVcs.CLcs.IRarXiv:1411.5726v22014End-to-End Dense Video Captioning with Masked Transformer
Luowei Zhou, Yingbo Zhou, Jason J. Corso +2
cs.CVarXiv:1804.00819v12018End-to-End Comparative Attention Networks for Person Re-identification
Hao Liu, Jiashi Feng, Meibin Qi +2
cs.CVarXiv:1606.04404v22016Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation
Weizhi Tao, Zengwang Jin, Xiao Wang +1
cs.CVcs.ROarXiv:2608.24366v12026Video Salient Object Detection via Fully Convolutional Networks
Wenguan Wang, Jianbing Shen, Ling Shao
cs.CVarXiv:1702.00871v32017StyleFlow: Attribute-conditioned Exploration of StyleGAN-Generated Images using Conditional Continuous Normalizing Flows
Rameen Abdal, Peihao Zhu, Niloy Mitra +1
cs.CVcs.GRarXiv:2008.02401v22020PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection
Shaoshuai Shi, Li Jiang, Jiajun Deng +5
cs.CVarXiv:2102.00463v32021Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision
Xinchen Yan, Jimei Yang, Ersin Yumer +2
cs.CVcs.GRcs.LGarXiv:1612.00814v32016WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation
Jongheon Jeong, Yang Zou, Taewan Kim +3
cs.CVcs.AIcs.CLarXiv:2303.14814v12023PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment
Ziqi Cui, Shangyu Lou
cs.CVcs.AIcs.IRarXiv:2608.24133v12026AANet: Adaptive Aggregation Network for Efficient Stereo Matching
Haofei Xu, Juyong Zhang
cs.CVarXiv:2004.09548v12020Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
Patrick Schramowski, Manuel Brack, Björn Deiseroth +1
cs.CVcs.AIcs.LGarXiv:2211.05105v42022LeFlow: Generative Latent Flow Planning for World Models
Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu +1
cs.CVarXiv:2608.24855v12026Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting
Zeyu Yang, Hongye Yang, Zijie Pan +1
cs.CVarXiv:2310.10642v32023MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang +7
cs.LGcs.AIcs.CLarXiv:2206.08853v22022It Is Not the Journey but the Destination: Endpoint Conditioned Trajectory Prediction
Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal +4
cs.CVcs.LGarXiv:2004.02025v32020Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman +2
cs.AIcs.CLcs.CVarXiv:2310.04406v32023Sliced Wasserstein Discrepancy for Unsupervised Domain Adaptation
Chen-Yu Lee, Tanmay Batra, Mohammad Haris Baig +1
cs.CVcs.LGstat.MLarXiv:1903.04064v12019Finding Action Tubes
Georgia Gkioxari, Jitendra Malik
cs.CVarXiv:1411.6031v12014A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
Jian Liang, Ran He, Tieniu Tan
cs.LGcs.AIcs.CVarXiv:2303.15361v22023Learning to Reason: End-to-End Module Networks for Visual Question Answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach +2
cs.CVarXiv:1704.05526v32017Spatial-Phase Shallow Learning: Rethinking Face Forgery Detection in Frequency Domain
Honggu Liu, Xiaodan Li, Wenbo Zhou +5
cs.CVarXiv:2103.01856v32021VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Haoxin Chen, Menghan Xia, Yingqing He +9
cs.CVarXiv:2310.19512v12023Camera Style Adaptation for Person Re-identification
Zhun Zhong, Liang Zheng, Zhedong Zheng +2
cs.CVarXiv:1711.10295v22017Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
Xingyi Zhou, Qixing Huang, Xiao Sun +2
cs.CVarXiv:1704.02447v22017A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
Xiaolong Wang, Abhinav Shrivastava, Abhinav Gupta
cs.CVarXiv:1704.03414v12017GPT-4V(ision) is a Generalist Web Agent, if Grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil +2
cs.IRcs.AIcs.CLarXiv:2401.01614v22024OneFormer: One Transformer to Rule Universal Image Segmentation
Jitesh Jain, Jiachen Li, MangTik Chiu +3
cs.CVarXiv:2211.06220v22022SWAD: Domain Generalization by Seeking Flat Minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee +4
cs.LGcs.CVarXiv:2102.08604v42021Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
Haoning Wu, Zicheng Zhang, Weixia Zhang +11
cs.CVcs.CLcs.LGarXiv:2312.17090v12023Tangent Convolutions for Dense Prediction in 3D
Maxim Tatarchenko, Jaesik Park, Vladlen Koltun +1
cs.CVarXiv:1807.02443v12018Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG
Zhe Jin, Zhimin Lin, Bin Zheng +2
cs.CVcs.AIarXiv:2608.23011v12026MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
Liangtao Shi, Jinxia Xie, Xiantao Hu +1
cs.CVarXiv:2608.23234v12026A Twofold Siamese Network for Real-Time Object Tracking
Anfeng He, Chong Luo, Xinmei Tian +1
cs.CVarXiv:1802.08817v12018Region Proposal by Guided Anchoring
Jiaqi Wang, Kai Chen, Shuo Yang +2
cs.CVarXiv:1901.03278v22019Learning 3D Human Dynamics from Video
Angjoo Kanazawa, Jason Y. Zhang, Panna Felsen +1
cs.CVarXiv:1812.01601v42018Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026
Jinxing Zhou, Suiyi Zhao, Yanghao Zhou +1
cs.MMcs.CVcs.SDarXiv:2608.22337v12026Incremental Learning of Object Detectors without Catastrophic Forgetting
Konstantin Shmelkov, Cordelia Schmid, Karteek Alahari
cs.CVarXiv:1708.06977v12017StegaStamp: Invisible Hyperlinks in Physical Photographs
Matthew Tancik, Ben Mildenhall, Ren Ng
cs.CVarXiv:1904.05343v22019Platonic Representation Hypothesis on World Models
Wenhow Li, Chengwei MA, Hui Xiong +2
cs.CVarXiv:2608.23720v12026Learning Data Augmentation Strategies for Object Detection
Barret Zoph, Ekin D. Cubuk, Golnaz Ghiasi +3
cs.CVcs.LGarXiv:1906.11172v12019Deep Global Registration
Christopher Choy, Wei Dong, Vladlen Koltun
cs.CVcs.CGcs.LGarXiv:2004.11540v22020Generalizing Face Forgery Detection with High-frequency Features
Yuchen Luo, Yong Zhang, Junchi Yan +1
cs.CVarXiv:2103.12376v12021MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
Huiyu Wang, Yukun Zhu, Hartwig Adam +2
cs.CVarXiv:2012.00759v32020Cross-Generation Optimization of YOLOv26, YOLOv11, and YOLOv8 for Fine-Grained Small-Object Detection and Instance Segmentation in Complex Orchards
Ranjan Sapkota, Manoj Karkee
cs.CVarXiv:2608.23636v12026Large-scale Multi-view Subspace Clustering in Linear Time
Zhao Kang, Wangtao Zhou, Zhitong Zhao +3
cs.LGcs.CVstat.MLarXiv:1911.09290v12019PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation
Yang Zhang, Zixiang Zhou, Philip David +4
cs.CVarXiv:2003.14032v22020Accelerating Eulerian Fluid Simulation With Convolutional Networks
Jonathan Tompson, Kristofer Schlachter, Pablo Sprechmann +1
cs.CVarXiv:1607.03597v72016Meta R-CNN : Towards General Solver for Instance-level Few-shot Learning
Xiaopeng Yan, Ziliang Chen, Anni Xu +3
cs.CVcs.LGarXiv:1909.13032v22019Invariance Matters: Exemplar Memory for Domain Adaptive Person Re-identification
Zhun Zhong, Liang Zheng, Zhiming Luo +2
cs.CVcs.LGarXiv:1904.01990v12019Open Set Domain Adaptation by Backpropagation
Kuniaki Saito, Shohei Yamamoto, Yoshitaka Ushiku +1
cs.CVarXiv:1804.10427v22018Low-Light Image Enhancement with Normalizing Flow
Yufei Wang, Renjie Wan, Wenhan Yang +3
eess.IVcs.CVarXiv:2109.05923v12021Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language
Songyang Zhang, Houwen Peng, Jianlong Fu +1
cs.CVcs.IRcs.MMarXiv:1912.03590v32019Visual Semantic Reasoning for Image-Text Matching
Kunpeng Li, Yulun Zhang, Kai Li +2
cs.CVarXiv:1909.02701v12019Learning to Compose Dynamic Tree Structures for Visual Contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu +2
cs.CVarXiv:1812.01880v12018$A^2$-Nets: Double Attention Networks
Yunpeng Chen, Yannis Kalantidis, Jianshu Li +2
cs.CVarXiv:1810.11579v12018A study of the effect of JPG compression on adversarial images
Gintare Karolina Dziugaite, Zoubin Ghahramani, Daniel M. Roy
cs.CVcs.LGarXiv:1608.00853v12016Asymmetric Tri-training for Unsupervised Domain Adaptation
Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada
cs.CVcs.AIarXiv:1702.08400v32017