Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,641 to 2,700 of 18,785
Effects of Degradations on Deep Neural Network Architectures
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya +1
cs.CVeess.IVarXiv:1807.10108v62018Adapting Mask-RCNN for Automatic Nucleus Segmentation
Jeremiah W. Johnson
cs.CVcs.LGarXiv:1805.00500v12018Model Watermarking for Image Processing Networks
Jie Zhang, Dongdong Chen, Jing Liao +5
cs.MMcs.CVeess.IVarXiv:2002.11088v12020GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani
cs.CVcs.AIcs.CLarXiv:2207.09666v12022Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?
Wenhao Wu, Haipeng Luo, Bo Fang +2
cs.CVarXiv:2301.00184v32022InstaGAN: Instance-aware Image-to-Image Translation
Sangwoo Mo, Minsu Cho, Jinwoo Shin
cs.LGcs.CVstat.MLarXiv:1812.10889v22018Mask Transfiner for High-Quality Instance Segmentation
Lei Ke, Martin Danelljan, Xia Li +3
cs.CVarXiv:2111.13673v12021PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds
Jinyu Li, Chenxu Luo, Xiaodong Yang
cs.CVarXiv:2305.04925v12023SiT: Self-supervised vIsion Transformer
Sara Atito, Muhammad Awais, Josef Kittler
cs.CVcs.LGarXiv:2104.03602v32021Scaling Up Influence Functions
Andrea Schioppa, Polina Zablotskaia, David Vilar +1
cs.LGcs.CLcs.CVarXiv:2112.03052v12021ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact
Samarth Brahmbhatt, Ankur Handa, James Hays +1
cs.ROcs.CVarXiv:1904.03754v32019WonderJourney: Going from Anywhere to Everywhere
Hong-Xing Yu, Haoyi Duan, Junhwa Hur +8
cs.CVcs.GRarXiv:2312.03884v22023Mapping the world population one building at a time
Tobias G. Tiecke, Xianming Liu, Amy Zhang +8
cs.CVarXiv:1712.05839v12017Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
Fan Bao, Chendong Xiang, Gang Yue +7
cs.CVcs.LGarXiv:2405.04233v12024BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models
Bowen Wen, Kostas Bekris
cs.CVcs.AIcs.GRarXiv:2108.00516v12021SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks
Xiang Xu, Karl D. D. Willis, Joseph G. Lambourne +3
cs.CVcs.LGarXiv:2207.04632v12022Leaf Counting with Deep Convolutional and Deconvolutional Networks
Shubhra Aich, Ian Stavness
cs.CVarXiv:1708.07570v22017F-formation Detection: Individuating Free-standing Conversational Groups in Images
Francesco Setti, Chris Russell, Chiara Bassetti +1
cs.CVarXiv:1409.2702v12014DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models
Jamie Wynn, Daniyar Turmukhambetov
cs.CVarXiv:2302.12231v32023RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank
Quentin Garrido, Randall Balestriero, Laurent Najman +1
cs.LGcs.AIcs.CVarXiv:2210.02885v32022A Study of Face Obfuscation in ImageNet
Kaiyu Yang, Jacqueline Yau, Li Fei-Fei +2
cs.CVarXiv:2103.06191v32021ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Peng Wang, Shijie Wang, Junyang Lin +5
cs.CVcs.CLcs.SDarXiv:2305.11172v12023ShuffleMixer: An Efficient ConvNet for Image Super-Resolution
Long Sun, Jinshan Pan, Jinhui Tang
cs.CVarXiv:2205.15175v12022MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining
Di Wang, Jing Zhang, Minqiang Xu +8
cs.CVarXiv:2403.13430v22024SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
Yuxuan Zhang, Yiren Song, Jiaming Liu +8
cs.CVarXiv:2312.16272v22023Learning to Continually Learn
Shawn Beaulieu, Lapo Frati, Thomas Miconi +4
cs.LGcs.CVcs.NEarXiv:2002.09571v22020A Realistic Dataset and Baseline Temporal Model for Early Drowsiness Detection
Reza Ghoddoosian, Marnim Galib, Vassilis Athitsos
cs.CVarXiv:1904.07312v12019One-Step Image Translation with Text-to-Image Models
Gaurav Parmar, Taesung Park, Srinivasa Narasimhan +1
cs.CVcs.GRcs.LGarXiv:2403.12036v12024RAMP-CNN: A Novel Neural Network for Enhanced Automotive Radar Object Recognition
Xiangyu Gao, Guanbin Xing, Sumit Roy +1
eess.SPcs.AIcs.CVarXiv:2011.08981v22020Re-thinking Co-Salient Object Detection
Deng-Ping Fan, Tengpeng Li, Zheng Lin +5
cs.CVarXiv:2007.03380v42020Hierarchical Integration Diffusion Model for Realistic Image Deblurring
Zheng Chen, Yulun Zhang, Ding Liu +4
cs.CVarXiv:2305.12966v42023Natural Language Descriptions of Deep Visual Features
Evan Hernandez, Sarah Schwettmann, David Bau +3
cs.CVcs.AIcs.CLarXiv:2201.11114v22022TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation
Dongxu Li, Chenchen Xu, Xin Yu +4
cs.CVcs.AIcs.HCarXiv:2010.05468v12020ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Yucheng Han, Chi Zhang, Xin Chen +5
cs.CVcs.CLarXiv:2311.16483v12023Pose-Invariant 3D Face Alignment
Amin Jourabloo, Xiaoming Liu
cs.CVarXiv:1506.03799v12015Can I Trust Your Answer? Visually Grounded Video Question Answering
Junbin Xiao, Angela Yao, Yicong Li +1
cs.CVcs.AIcs.MMarXiv:2309.01327v22023Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models
Rohit Gandikota, Joanna Materzynska, Tingrui Zhou +2
cs.CVarXiv:2311.12092v22023Operation-Aware Soft Channel Pruning using Differentiable Masks
Minsoo Kang, Bohyung Han
cs.LGcs.CVstat.MLarXiv:2007.03938v22020On Success and Simplicity: A Second Look at Transferable Targeted Attacks
Zhengyu Zhao, Zhuoran Liu, Martha Larson
cs.LGcs.CRcs.CVarXiv:2012.11207v42020EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition
Gabriele Berton, Gabriele Trivigno, Barbara Caputo +1
cs.CVarXiv:2308.10832v12023Editing Text in the Wild
Liang Wu, Chengquan Zhang, Jiaming Liu +4
cs.CVarXiv:1908.03047v12019Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer
Frederic Z. Zhang, Dylan Campbell, Stephen Gould
cs.CVcs.AIcs.LGarXiv:2112.01838v22021Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction
Yingxue Xu, Hao Chen
cs.CVarXiv:2306.08330v22023Crowdsourcing in Computer Vision
Adriana Kovashka, Olga Russakovsky, Li Fei-Fei +1
cs.CVcs.HCarXiv:1611.02145v12016Affective Image Content Analysis: Two Decades Review and New Perspectives
Sicheng Zhao, Xingxu Yao, Jufeng Yang +5
cs.CVcs.AIcs.MMarXiv:2106.16125v12021Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
Si Liu, Xiaodan Liang, Luoqi Liu +6
cs.CVarXiv:1504.01220v12015Breaking Darknet CAPTCHAs with general purpose LLM
Benjamin Fehrensen, Jens Hubler
cs.CRcs.CVarXiv:2608.28794v12026Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
Xuechao Zou, Yi Zhou, Kai Li +4
cs.CVcs.AIarXiv:2609.07670v12026Towards Universal Representation Learning for Deep Face Recognition
Yichun Shi, Xiang Yu, Kihyuk Sohn +2
cs.CVarXiv:2002.11841v12020Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6
cs.CVcs.LGarXiv:2609.08084v12026LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning
Atsuyuki Miyai, Qing Yu, Go Irie +1
cs.CVarXiv:2306.01293v32023SC^2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration
Zhi Chen, Kun Sun, Fan Yang +1
cs.CVarXiv:2203.14453v12022OadTR: Online Action Detection with Transformers
Xiang Wang, Shiwei Zhang, Zhiwu Qing +4
cs.CVarXiv:2106.11149v12021Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
Jaemin Cho, Yushi Hu, Roopal Garg +6
cs.CVcs.AIcs.CLarXiv:2310.18235v42023J$\hat{\text{A}}$A-Net: Joint Facial Action Unit Detection and Face Alignment via Adaptive Attention
Zhiwen Shao, Zhilei Liu, Jianfei Cai +1
cs.CVarXiv:2003.08834v32020A survey of advances in vision-based vehicle re-identification
Sultan Daud Khan, Habib Ullah
cs.CVcs.AIarXiv:1905.13258v12019Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
Hamza Kheddar
cs.CRcs.AIcs.CLarXiv:2408.07583v22024RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
Hejun Wang, Jinxi Li, Junwei Jiang +4
cs.CVcs.AIcs.GRarXiv:2609.07414v12026DuPLO: A DUal view Point deep Learning architecture for time series classificatiOn
Roberto Interdonato, Dino Ienco, Raffaele Gaetano +1
cs.CVarXiv:1809.07589v12018ktrain: A Low-Code Library for Augmented Machine Learning
Arun S. Maiya
cs.LGcs.CLcs.CVarXiv:2004.10703v52020