Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,641 to 2,700 of 18,785

  1. Effects of Degradations on Deep Neural Network Architectures

    Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya +1

    cs.CVeess.IVarXiv:1807.10108v62018
  2. Adapting Mask-RCNN for Automatic Nucleus Segmentation

    Jeremiah W. Johnson

    cs.CVcs.LGarXiv:1805.00500v12018
  3. Model Watermarking for Image Processing Networks

    Jie Zhang, Dongdong Chen, Jing Liao +5

    cs.MMcs.CVeess.IVarXiv:2002.11088v12020
  4. GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features

    Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani

    cs.CVcs.AIcs.CLarXiv:2207.09666v12022
  5. Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?

    Wenhao Wu, Haipeng Luo, Bo Fang +2

    cs.CVarXiv:2301.00184v32022
  6. InstaGAN: Instance-aware Image-to-Image Translation

    Sangwoo Mo, Minsu Cho, Jinwoo Shin

    cs.LGcs.CVstat.MLarXiv:1812.10889v22018
  7. Mask Transfiner for High-Quality Instance Segmentation

    Lei Ke, Martin Danelljan, Xia Li +3

    cs.CVarXiv:2111.13673v12021
  8. PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds

    Jinyu Li, Chenxu Luo, Xiaodong Yang

    cs.CVarXiv:2305.04925v12023
  9. SiT: Self-supervised vIsion Transformer

    Sara Atito, Muhammad Awais, Josef Kittler

    cs.CVcs.LGarXiv:2104.03602v32021
  10. Scaling Up Influence Functions

    Andrea Schioppa, Polina Zablotskaia, David Vilar +1

    cs.LGcs.CLcs.CVarXiv:2112.03052v12021
  11. ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact

    Samarth Brahmbhatt, Ankur Handa, James Hays +1

    cs.ROcs.CVarXiv:1904.03754v32019
  12. WonderJourney: Going from Anywhere to Everywhere

    Hong-Xing Yu, Haoyi Duan, Junhwa Hur +8

    cs.CVcs.GRarXiv:2312.03884v22023
  13. Mapping the world population one building at a time

    Tobias G. Tiecke, Xianming Liu, Amy Zhang +8

    cs.CVarXiv:1712.05839v12017
  14. Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

    Fan Bao, Chendong Xiang, Gang Yue +7

    cs.CVcs.LGarXiv:2405.04233v12024
  15. BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models

    Bowen Wen, Kostas Bekris

    cs.CVcs.AIcs.GRarXiv:2108.00516v12021
  16. SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks

    Xiang Xu, Karl D. D. Willis, Joseph G. Lambourne +3

    cs.CVcs.LGarXiv:2207.04632v12022
  17. Leaf Counting with Deep Convolutional and Deconvolutional Networks

    Shubhra Aich, Ian Stavness

    cs.CVarXiv:1708.07570v22017
  18. F-formation Detection: Individuating Free-standing Conversational Groups in Images

    Francesco Setti, Chris Russell, Chiara Bassetti +1

    cs.CVarXiv:1409.2702v12014
  19. DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models

    Jamie Wynn, Daniyar Turmukhambetov

    cs.CVarXiv:2302.12231v32023
  20. RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank

    Quentin Garrido, Randall Balestriero, Laurent Najman +1

    cs.LGcs.AIcs.CVarXiv:2210.02885v32022
  21. A Study of Face Obfuscation in ImageNet

    Kaiyu Yang, Jacqueline Yau, Li Fei-Fei +2

    cs.CVarXiv:2103.06191v32021
  22. ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    Peng Wang, Shijie Wang, Junyang Lin +5

    cs.CVcs.CLcs.SDarXiv:2305.11172v12023
  23. ShuffleMixer: An Efficient ConvNet for Image Super-Resolution

    Long Sun, Jinshan Pan, Jinhui Tang

    cs.CVarXiv:2205.15175v12022
  24. MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining

    Di Wang, Jing Zhang, Minqiang Xu +8

    cs.CVarXiv:2403.13430v22024
  25. SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation

    Yuxuan Zhang, Yiren Song, Jiaming Liu +8

    cs.CVarXiv:2312.16272v22023
  26. Learning to Continually Learn

    Shawn Beaulieu, Lapo Frati, Thomas Miconi +4

    cs.LGcs.CVcs.NEarXiv:2002.09571v22020
  27. A Realistic Dataset and Baseline Temporal Model for Early Drowsiness Detection

    Reza Ghoddoosian, Marnim Galib, Vassilis Athitsos

    cs.CVarXiv:1904.07312v12019
  28. One-Step Image Translation with Text-to-Image Models

    Gaurav Parmar, Taesung Park, Srinivasa Narasimhan +1

    cs.CVcs.GRcs.LGarXiv:2403.12036v12024
  29. RAMP-CNN: A Novel Neural Network for Enhanced Automotive Radar Object Recognition

    Xiangyu Gao, Guanbin Xing, Sumit Roy +1

    eess.SPcs.AIcs.CVarXiv:2011.08981v22020
  30. Re-thinking Co-Salient Object Detection

    Deng-Ping Fan, Tengpeng Li, Zheng Lin +5

    cs.CVarXiv:2007.03380v42020
  31. Hierarchical Integration Diffusion Model for Realistic Image Deblurring

    Zheng Chen, Yulun Zhang, Ding Liu +4

    cs.CVarXiv:2305.12966v42023
  32. Natural Language Descriptions of Deep Visual Features

    Evan Hernandez, Sarah Schwettmann, David Bau +3

    cs.CVcs.AIcs.CLarXiv:2201.11114v22022
  33. TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation

    Dongxu Li, Chenchen Xu, Xin Yu +4

    cs.CVcs.AIcs.HCarXiv:2010.05468v12020
  34. ChartLlama: A Multimodal LLM for Chart Understanding and Generation

    Yucheng Han, Chi Zhang, Xin Chen +5

    cs.CVcs.CLarXiv:2311.16483v12023
  35. Pose-Invariant 3D Face Alignment

    Amin Jourabloo, Xiaoming Liu

    cs.CVarXiv:1506.03799v12015
  36. Can I Trust Your Answer? Visually Grounded Video Question Answering

    Junbin Xiao, Angela Yao, Yicong Li +1

    cs.CVcs.AIcs.MMarXiv:2309.01327v22023
  37. Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models

    Rohit Gandikota, Joanna Materzynska, Tingrui Zhou +2

    cs.CVarXiv:2311.12092v22023
  38. Operation-Aware Soft Channel Pruning using Differentiable Masks

    Minsoo Kang, Bohyung Han

    cs.LGcs.CVstat.MLarXiv:2007.03938v22020
  39. On Success and Simplicity: A Second Look at Transferable Targeted Attacks

    Zhengyu Zhao, Zhuoran Liu, Martha Larson

    cs.LGcs.CRcs.CVarXiv:2012.11207v42020
  40. EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition

    Gabriele Berton, Gabriele Trivigno, Barbara Caputo +1

    cs.CVarXiv:2308.10832v12023
  41. Editing Text in the Wild

    Liang Wu, Chengquan Zhang, Jiaming Liu +4

    cs.CVarXiv:1908.03047v12019
  42. Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer

    Frederic Z. Zhang, Dylan Campbell, Stephen Gould

    cs.CVcs.AIcs.LGarXiv:2112.01838v22021
  43. Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction

    Yingxue Xu, Hao Chen

    cs.CVarXiv:2306.08330v22023
  44. Crowdsourcing in Computer Vision

    Adriana Kovashka, Olga Russakovsky, Li Fei-Fei +1

    cs.CVcs.HCarXiv:1611.02145v12016
  45. Affective Image Content Analysis: Two Decades Review and New Perspectives

    Sicheng Zhao, Xingxu Yao, Jufeng Yang +5

    cs.CVcs.AIcs.MMarXiv:2106.16125v12021
  46. Matching-CNN Meets KNN: Quasi-Parametric Human Parsing

    Si Liu, Xiaodan Liang, Luoqi Liu +6

    cs.CVarXiv:1504.01220v12015
  47. Breaking Darknet CAPTCHAs with general purpose LLM

    Benjamin Fehrensen, Jens Hubler

    cs.CRcs.CVarXiv:2608.28794v12026
  48. Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

    Xuechao Zou, Yi Zhou, Kai Li +4

    cs.CVcs.AIarXiv:2609.07670v12026
  49. Towards Universal Representation Learning for Deep Face Recognition

    Yichun Shi, Xiang Yu, Kihyuk Sohn +2

    cs.CVarXiv:2002.11841v12020
  50. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

    Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6

    cs.CVcs.LGarXiv:2609.08084v12026
  51. LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning

    Atsuyuki Miyai, Qing Yu, Go Irie +1

    cs.CVarXiv:2306.01293v32023
  52. SC^2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration

    Zhi Chen, Kun Sun, Fan Yang +1

    cs.CVarXiv:2203.14453v12022
  53. OadTR: Online Action Detection with Transformers

    Xiang Wang, Shiwei Zhang, Zhiwu Qing +4

    cs.CVarXiv:2106.11149v12021
  54. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

    Jaemin Cho, Yushi Hu, Roopal Garg +6

    cs.CVcs.AIcs.CLarXiv:2310.18235v42023
  55. J$\hat{\text{A}}$A-Net: Joint Facial Action Unit Detection and Face Alignment via Adaptive Attention

    Zhiwen Shao, Zhilei Liu, Jianfei Cai +1

    cs.CVarXiv:2003.08834v32020
  56. A survey of advances in vision-based vehicle re-identification

    Sultan Daud Khan, Habib Ullah

    cs.CVcs.AIarXiv:1905.13258v12019
  57. Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey

    Hamza Kheddar

    cs.CRcs.AIcs.CLarXiv:2408.07583v22024
  58. RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

    Hejun Wang, Jinxi Li, Junwei Jiang +4

    cs.CVcs.AIcs.GRarXiv:2609.07414v12026
  59. DuPLO: A DUal view Point deep Learning architecture for time series classificatiOn

    Roberto Interdonato, Dino Ienco, Raffaele Gaetano +1

    cs.CVarXiv:1809.07589v12018
  60. ktrain: A Low-Code Library for Augmented Machine Learning

    Arun S. Maiya

    cs.LGcs.CLcs.CVarXiv:2004.10703v52020