Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,701 to 2,760 of 18,855
Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification
Xudong Tian, Zhizhong Zhang, Shaohui Lin +3
cs.CVarXiv:2104.02862v22021Continual Learning of a Mixed Sequence of Similar and Dissimilar Tasks
Zixuan Ke, Bing Liu, Xingchang Huang
cs.LGcs.AIcs.CVarXiv:2112.10017v12021Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective
Yongjin Cui, Xiaohui Fan
cs.AIcs.CVarXiv:2609.07213v12026Tiny SSD: A Tiny Single-shot Detection Deep Convolutional Neural Network for Real-time Embedded Object Detection
Alexander Wong, Mohammad Javad Shafiee, Francis Li +1
cs.CVcs.AIcs.NEarXiv:1802.06488v12018AFDetV2: Rethinking the Necessity of the Second Stage for Object Detection from Point Clouds
Yihan Hu, Zhuangzhuang Ding, Runzhou Ge +4
cs.CVarXiv:2112.09205v22021Learning Transferable Adversarial Examples via Ghost Networks
Yingwei Li, Song Bai, Yuyin Zhou +3
cs.CVcs.LGarXiv:1812.03413v32018Infinigen Indoors: Photorealistic Indoor Scenes using Procedural Generation
Alexander Raistrick, Lingjie Mei, Karhan Kayan +9
cs.CVarXiv:2406.11824v12024Bags of Local Convolutional Features for Scalable Instance Search
Eva Mohedano, Amaia Salvador, Kevin McGuinness +3
cs.CVcs.MMarXiv:1604.04653v12016Recurrently Exploring Class-wise Attention in A Hybrid Convolutional and Bidirectional LSTM Network for Multi-label Aerial Image Classification
Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu
cs.CVarXiv:1807.11245v22018Look, Listen, and Act: Towards Audio-Visual Embodied Navigation
Chuang Gan, Yiwei Zhang, Jiajun Wu +2
cs.CVcs.LGcs.ROarXiv:1912.11684v22019Effects of Degradations on Deep Neural Network Architectures
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya +1
cs.CVeess.IVarXiv:1807.10108v62018Adapting Mask-RCNN for Automatic Nucleus Segmentation
Jeremiah W. Johnson
cs.CVcs.LGarXiv:1805.00500v12018Model Watermarking for Image Processing Networks
Jie Zhang, Dongdong Chen, Jing Liao +5
cs.MMcs.CVeess.IVarXiv:2002.11088v12020GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani
cs.CVcs.AIcs.CLarXiv:2207.09666v12022Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?
Wenhao Wu, Haipeng Luo, Bo Fang +2
cs.CVarXiv:2301.00184v32022InstaGAN: Instance-aware Image-to-Image Translation
Sangwoo Mo, Minsu Cho, Jinwoo Shin
cs.LGcs.CVstat.MLarXiv:1812.10889v22018Mask Transfiner for High-Quality Instance Segmentation
Lei Ke, Martin Danelljan, Xia Li +3
cs.CVarXiv:2111.13673v12021PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds
Jinyu Li, Chenxu Luo, Xiaodong Yang
cs.CVarXiv:2305.04925v12023SiT: Self-supervised vIsion Transformer
Sara Atito, Muhammad Awais, Josef Kittler
cs.CVcs.LGarXiv:2104.03602v32021Scaling Up Influence Functions
Andrea Schioppa, Polina Zablotskaia, David Vilar +1
cs.LGcs.CLcs.CVarXiv:2112.03052v12021ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact
Samarth Brahmbhatt, Ankur Handa, James Hays +1
cs.ROcs.CVarXiv:1904.03754v32019WonderJourney: Going from Anywhere to Everywhere
Hong-Xing Yu, Haoyi Duan, Junhwa Hur +8
cs.CVcs.GRarXiv:2312.03884v22023Mapping the world population one building at a time
Tobias G. Tiecke, Xianming Liu, Amy Zhang +8
cs.CVarXiv:1712.05839v12017Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
Fan Bao, Chendong Xiang, Gang Yue +7
cs.CVcs.LGarXiv:2405.04233v12024BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models
Bowen Wen, Kostas Bekris
cs.CVcs.AIcs.GRarXiv:2108.00516v12021SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks
Xiang Xu, Karl D. D. Willis, Joseph G. Lambourne +3
cs.CVcs.LGarXiv:2207.04632v12022Leaf Counting with Deep Convolutional and Deconvolutional Networks
Shubhra Aich, Ian Stavness
cs.CVarXiv:1708.07570v22017F-formation Detection: Individuating Free-standing Conversational Groups in Images
Francesco Setti, Chris Russell, Chiara Bassetti +1
cs.CVarXiv:1409.2702v12014DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models
Jamie Wynn, Daniyar Turmukhambetov
cs.CVarXiv:2302.12231v32023RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank
Quentin Garrido, Randall Balestriero, Laurent Najman +1
cs.LGcs.AIcs.CVarXiv:2210.02885v32022A Study of Face Obfuscation in ImageNet
Kaiyu Yang, Jacqueline Yau, Li Fei-Fei +2
cs.CVarXiv:2103.06191v32021ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Peng Wang, Shijie Wang, Junyang Lin +5
cs.CVcs.CLcs.SDarXiv:2305.11172v12023ShuffleMixer: An Efficient ConvNet for Image Super-Resolution
Long Sun, Jinshan Pan, Jinhui Tang
cs.CVarXiv:2205.15175v12022MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining
Di Wang, Jing Zhang, Minqiang Xu +8
cs.CVarXiv:2403.13430v22024SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
Yuxuan Zhang, Yiren Song, Jiaming Liu +8
cs.CVarXiv:2312.16272v22023Learning to Continually Learn
Shawn Beaulieu, Lapo Frati, Thomas Miconi +4
cs.LGcs.CVcs.NEarXiv:2002.09571v22020A Realistic Dataset and Baseline Temporal Model for Early Drowsiness Detection
Reza Ghoddoosian, Marnim Galib, Vassilis Athitsos
cs.CVarXiv:1904.07312v12019One-Step Image Translation with Text-to-Image Models
Gaurav Parmar, Taesung Park, Srinivasa Narasimhan +1
cs.CVcs.GRcs.LGarXiv:2403.12036v12024RAMP-CNN: A Novel Neural Network for Enhanced Automotive Radar Object Recognition
Xiangyu Gao, Guanbin Xing, Sumit Roy +1
eess.SPcs.AIcs.CVarXiv:2011.08981v22020Re-thinking Co-Salient Object Detection
Deng-Ping Fan, Tengpeng Li, Zheng Lin +5
cs.CVarXiv:2007.03380v42020Hierarchical Integration Diffusion Model for Realistic Image Deblurring
Zheng Chen, Yulun Zhang, Ding Liu +4
cs.CVarXiv:2305.12966v42023Natural Language Descriptions of Deep Visual Features
Evan Hernandez, Sarah Schwettmann, David Bau +3
cs.CVcs.AIcs.CLarXiv:2201.11114v22022TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation
Dongxu Li, Chenchen Xu, Xin Yu +4
cs.CVcs.AIcs.HCarXiv:2010.05468v12020ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Yucheng Han, Chi Zhang, Xin Chen +5
cs.CVcs.CLarXiv:2311.16483v12023Pose-Invariant 3D Face Alignment
Amin Jourabloo, Xiaoming Liu
cs.CVarXiv:1506.03799v12015Can I Trust Your Answer? Visually Grounded Video Question Answering
Junbin Xiao, Angela Yao, Yicong Li +1
cs.CVcs.AIcs.MMarXiv:2309.01327v22023Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models
Rohit Gandikota, Joanna Materzynska, Tingrui Zhou +2
cs.CVarXiv:2311.12092v22023Operation-Aware Soft Channel Pruning using Differentiable Masks
Minsoo Kang, Bohyung Han
cs.LGcs.CVstat.MLarXiv:2007.03938v22020On Success and Simplicity: A Second Look at Transferable Targeted Attacks
Zhengyu Zhao, Zhuoran Liu, Martha Larson
cs.LGcs.CRcs.CVarXiv:2012.11207v42020EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition
Gabriele Berton, Gabriele Trivigno, Barbara Caputo +1
cs.CVarXiv:2308.10832v12023Editing Text in the Wild
Liang Wu, Chengquan Zhang, Jiaming Liu +4
cs.CVarXiv:1908.03047v12019Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer
Frederic Z. Zhang, Dylan Campbell, Stephen Gould
cs.CVcs.AIcs.LGarXiv:2112.01838v22021Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction
Yingxue Xu, Hao Chen
cs.CVarXiv:2306.08330v22023Crowdsourcing in Computer Vision
Adriana Kovashka, Olga Russakovsky, Li Fei-Fei +1
cs.CVcs.HCarXiv:1611.02145v12016Affective Image Content Analysis: Two Decades Review and New Perspectives
Sicheng Zhao, Xingxu Yao, Jufeng Yang +5
cs.CVcs.AIcs.MMarXiv:2106.16125v12021Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
Si Liu, Xiaodan Liang, Luoqi Liu +6
cs.CVarXiv:1504.01220v12015Breaking Darknet CAPTCHAs with general purpose LLM
Benjamin Fehrensen, Jens Hubler
cs.CRcs.CVarXiv:2608.28794v12026Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
Xuechao Zou, Yi Zhou, Kai Li +4
cs.CVcs.AIarXiv:2609.07670v12026Towards Universal Representation Learning for Deep Face Recognition
Yichun Shi, Xiang Yu, Kihyuk Sohn +2
cs.CVarXiv:2002.11841v12020Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6
cs.CVcs.LGarXiv:2609.08084v12026