Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,701 to 2,760 of 18,855

  1. Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification

    Xudong Tian, Zhizhong Zhang, Shaohui Lin +3

    cs.CVarXiv:2104.02862v22021
  2. Continual Learning of a Mixed Sequence of Similar and Dissimilar Tasks

    Zixuan Ke, Bing Liu, Xingchang Huang

    cs.LGcs.AIcs.CVarXiv:2112.10017v12021
  3. Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective

    Yongjin Cui, Xiaohui Fan

    cs.AIcs.CVarXiv:2609.07213v12026
  4. Tiny SSD: A Tiny Single-shot Detection Deep Convolutional Neural Network for Real-time Embedded Object Detection

    Alexander Wong, Mohammad Javad Shafiee, Francis Li +1

    cs.CVcs.AIcs.NEarXiv:1802.06488v12018
  5. AFDetV2: Rethinking the Necessity of the Second Stage for Object Detection from Point Clouds

    Yihan Hu, Zhuangzhuang Ding, Runzhou Ge +4

    cs.CVarXiv:2112.09205v22021
  6. Learning Transferable Adversarial Examples via Ghost Networks

    Yingwei Li, Song Bai, Yuyin Zhou +3

    cs.CVcs.LGarXiv:1812.03413v32018
  7. Infinigen Indoors: Photorealistic Indoor Scenes using Procedural Generation

    Alexander Raistrick, Lingjie Mei, Karhan Kayan +9

    cs.CVarXiv:2406.11824v12024
  8. Bags of Local Convolutional Features for Scalable Instance Search

    Eva Mohedano, Amaia Salvador, Kevin McGuinness +3

    cs.CVcs.MMarXiv:1604.04653v12016
  9. Recurrently Exploring Class-wise Attention in A Hybrid Convolutional and Bidirectional LSTM Network for Multi-label Aerial Image Classification

    Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu

    cs.CVarXiv:1807.11245v22018
  10. Look, Listen, and Act: Towards Audio-Visual Embodied Navigation

    Chuang Gan, Yiwei Zhang, Jiajun Wu +2

    cs.CVcs.LGcs.ROarXiv:1912.11684v22019
  11. Effects of Degradations on Deep Neural Network Architectures

    Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya +1

    cs.CVeess.IVarXiv:1807.10108v62018
  12. Adapting Mask-RCNN for Automatic Nucleus Segmentation

    Jeremiah W. Johnson

    cs.CVcs.LGarXiv:1805.00500v12018
  13. Model Watermarking for Image Processing Networks

    Jie Zhang, Dongdong Chen, Jing Liao +5

    cs.MMcs.CVeess.IVarXiv:2002.11088v12020
  14. GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features

    Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani

    cs.CVcs.AIcs.CLarXiv:2207.09666v12022
  15. Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?

    Wenhao Wu, Haipeng Luo, Bo Fang +2

    cs.CVarXiv:2301.00184v32022
  16. InstaGAN: Instance-aware Image-to-Image Translation

    Sangwoo Mo, Minsu Cho, Jinwoo Shin

    cs.LGcs.CVstat.MLarXiv:1812.10889v22018
  17. Mask Transfiner for High-Quality Instance Segmentation

    Lei Ke, Martin Danelljan, Xia Li +3

    cs.CVarXiv:2111.13673v12021
  18. PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds

    Jinyu Li, Chenxu Luo, Xiaodong Yang

    cs.CVarXiv:2305.04925v12023
  19. SiT: Self-supervised vIsion Transformer

    Sara Atito, Muhammad Awais, Josef Kittler

    cs.CVcs.LGarXiv:2104.03602v32021
  20. Scaling Up Influence Functions

    Andrea Schioppa, Polina Zablotskaia, David Vilar +1

    cs.LGcs.CLcs.CVarXiv:2112.03052v12021
  21. ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact

    Samarth Brahmbhatt, Ankur Handa, James Hays +1

    cs.ROcs.CVarXiv:1904.03754v32019
  22. WonderJourney: Going from Anywhere to Everywhere

    Hong-Xing Yu, Haoyi Duan, Junhwa Hur +8

    cs.CVcs.GRarXiv:2312.03884v22023
  23. Mapping the world population one building at a time

    Tobias G. Tiecke, Xianming Liu, Amy Zhang +8

    cs.CVarXiv:1712.05839v12017
  24. Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

    Fan Bao, Chendong Xiang, Gang Yue +7

    cs.CVcs.LGarXiv:2405.04233v12024
  25. BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models

    Bowen Wen, Kostas Bekris

    cs.CVcs.AIcs.GRarXiv:2108.00516v12021
  26. SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks

    Xiang Xu, Karl D. D. Willis, Joseph G. Lambourne +3

    cs.CVcs.LGarXiv:2207.04632v12022
  27. Leaf Counting with Deep Convolutional and Deconvolutional Networks

    Shubhra Aich, Ian Stavness

    cs.CVarXiv:1708.07570v22017
  28. F-formation Detection: Individuating Free-standing Conversational Groups in Images

    Francesco Setti, Chris Russell, Chiara Bassetti +1

    cs.CVarXiv:1409.2702v12014
  29. DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models

    Jamie Wynn, Daniyar Turmukhambetov

    cs.CVarXiv:2302.12231v32023
  30. RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank

    Quentin Garrido, Randall Balestriero, Laurent Najman +1

    cs.LGcs.AIcs.CVarXiv:2210.02885v32022
  31. A Study of Face Obfuscation in ImageNet

    Kaiyu Yang, Jacqueline Yau, Li Fei-Fei +2

    cs.CVarXiv:2103.06191v32021
  32. ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    Peng Wang, Shijie Wang, Junyang Lin +5

    cs.CVcs.CLcs.SDarXiv:2305.11172v12023
  33. ShuffleMixer: An Efficient ConvNet for Image Super-Resolution

    Long Sun, Jinshan Pan, Jinhui Tang

    cs.CVarXiv:2205.15175v12022
  34. MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining

    Di Wang, Jing Zhang, Minqiang Xu +8

    cs.CVarXiv:2403.13430v22024
  35. SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation

    Yuxuan Zhang, Yiren Song, Jiaming Liu +8

    cs.CVarXiv:2312.16272v22023
  36. Learning to Continually Learn

    Shawn Beaulieu, Lapo Frati, Thomas Miconi +4

    cs.LGcs.CVcs.NEarXiv:2002.09571v22020
  37. A Realistic Dataset and Baseline Temporal Model for Early Drowsiness Detection

    Reza Ghoddoosian, Marnim Galib, Vassilis Athitsos

    cs.CVarXiv:1904.07312v12019
  38. One-Step Image Translation with Text-to-Image Models

    Gaurav Parmar, Taesung Park, Srinivasa Narasimhan +1

    cs.CVcs.GRcs.LGarXiv:2403.12036v12024
  39. RAMP-CNN: A Novel Neural Network for Enhanced Automotive Radar Object Recognition

    Xiangyu Gao, Guanbin Xing, Sumit Roy +1

    eess.SPcs.AIcs.CVarXiv:2011.08981v22020
  40. Re-thinking Co-Salient Object Detection

    Deng-Ping Fan, Tengpeng Li, Zheng Lin +5

    cs.CVarXiv:2007.03380v42020
  41. Hierarchical Integration Diffusion Model for Realistic Image Deblurring

    Zheng Chen, Yulun Zhang, Ding Liu +4

    cs.CVarXiv:2305.12966v42023
  42. Natural Language Descriptions of Deep Visual Features

    Evan Hernandez, Sarah Schwettmann, David Bau +3

    cs.CVcs.AIcs.CLarXiv:2201.11114v22022
  43. TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation

    Dongxu Li, Chenchen Xu, Xin Yu +4

    cs.CVcs.AIcs.HCarXiv:2010.05468v12020
  44. ChartLlama: A Multimodal LLM for Chart Understanding and Generation

    Yucheng Han, Chi Zhang, Xin Chen +5

    cs.CVcs.CLarXiv:2311.16483v12023
  45. Pose-Invariant 3D Face Alignment

    Amin Jourabloo, Xiaoming Liu

    cs.CVarXiv:1506.03799v12015
  46. Can I Trust Your Answer? Visually Grounded Video Question Answering

    Junbin Xiao, Angela Yao, Yicong Li +1

    cs.CVcs.AIcs.MMarXiv:2309.01327v22023
  47. Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models

    Rohit Gandikota, Joanna Materzynska, Tingrui Zhou +2

    cs.CVarXiv:2311.12092v22023
  48. Operation-Aware Soft Channel Pruning using Differentiable Masks

    Minsoo Kang, Bohyung Han

    cs.LGcs.CVstat.MLarXiv:2007.03938v22020
  49. On Success and Simplicity: A Second Look at Transferable Targeted Attacks

    Zhengyu Zhao, Zhuoran Liu, Martha Larson

    cs.LGcs.CRcs.CVarXiv:2012.11207v42020
  50. EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition

    Gabriele Berton, Gabriele Trivigno, Barbara Caputo +1

    cs.CVarXiv:2308.10832v12023
  51. Editing Text in the Wild

    Liang Wu, Chengquan Zhang, Jiaming Liu +4

    cs.CVarXiv:1908.03047v12019
  52. Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer

    Frederic Z. Zhang, Dylan Campbell, Stephen Gould

    cs.CVcs.AIcs.LGarXiv:2112.01838v22021
  53. Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction

    Yingxue Xu, Hao Chen

    cs.CVarXiv:2306.08330v22023
  54. Crowdsourcing in Computer Vision

    Adriana Kovashka, Olga Russakovsky, Li Fei-Fei +1

    cs.CVcs.HCarXiv:1611.02145v12016
  55. Affective Image Content Analysis: Two Decades Review and New Perspectives

    Sicheng Zhao, Xingxu Yao, Jufeng Yang +5

    cs.CVcs.AIcs.MMarXiv:2106.16125v12021
  56. Matching-CNN Meets KNN: Quasi-Parametric Human Parsing

    Si Liu, Xiaodan Liang, Luoqi Liu +6

    cs.CVarXiv:1504.01220v12015
  57. Breaking Darknet CAPTCHAs with general purpose LLM

    Benjamin Fehrensen, Jens Hubler

    cs.CRcs.CVarXiv:2608.28794v12026
  58. Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

    Xuechao Zou, Yi Zhou, Kai Li +4

    cs.CVcs.AIarXiv:2609.07670v12026
  59. Towards Universal Representation Learning for Deep Face Recognition

    Yichun Shi, Xiang Yu, Kihyuk Sohn +2

    cs.CVarXiv:2002.11841v12020
  60. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

    Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6

    cs.CVcs.LGarXiv:2609.08084v12026