Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,941 to 9,000 of 18,866
Image Generation from Layout
Bo Zhao, Lili Meng, Weidong Yin +1
cs.CVeess.IVarXiv:1811.11389v32018MSeg: A Composite Dataset for Multi-domain Semantic Segmentation
John Lambert, Zhuang Liu, Ozan Sener +2
cs.CVarXiv:2112.13762v12021UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
Can Qin, Shu Zhang, Ning Yu +10
cs.CVcs.AIarXiv:2305.11147v32023A Transformer-Based Feature Segmentation and Region Alignment Method For UAV-View Geo-Localization
Ming Dai, Jianhong Hu, Jiedong Zhuang +1
cs.CVcs.AIarXiv:2201.09206v12022Video Propagation Networks
Varun Jampani, Raghudeep Gadde, Peter V. Gehler
cs.CVarXiv:1612.05478v32016Stabilizing Differentiable Architecture Search via Perturbation-based Regularization
Xiangning Chen, Cho-Jui Hsieh
cs.LGcs.CVstat.MLarXiv:2002.05283v32020Implicit Semantic Data Augmentation for Deep Networks
Yulin Wang, Xuran Pan, Shiji Song +3
cs.CVcs.LGstat.MLarXiv:1909.12220v52019Superpixel Segmentation with Fully Convolutional Networks
Fengting Yang, Qian Sun, Hailin Jin +1
cs.CVarXiv:2003.12929v12020MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology
Chaoyi Wu, Xiaoman Zhang, Ya Zhang +2
eess.IVcs.CLcs.CVarXiv:2301.02228v32023EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model
Xinya Ji, Hang Zhou, Kaisiyuan Wang +4
cs.CVarXiv:2205.15278v32022To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now
Yimeng Zhang, Jinghan Jia, Xin Chen +5
cs.CVarXiv:2310.11868v42023End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering
Youngjae Yu, Hyungjin Ko, Jongwook Choi +1
cs.CVarXiv:1610.02947v32016Deformable Shape Completion with Graph Convolutional Autoencoders
Or Litany, Alex Bronstein, Michael Bronstein +1
cs.CVarXiv:1712.00268v42017ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
Shaghayegh Kolli, Sina Emami, Moreno D'Incà +4
cs.CVcs.CLarXiv:2608.29847v12026Spatially-sparse convolutional neural networks
Benjamin Graham
cs.CVcs.NEarXiv:1409.6070v12014Deep Representation Learning on Long-tailed Data: A Learnable Embedding Augmentation Perspective
Jialun Liu, Yifan Sun, Chuchu Han +2
cs.CVarXiv:2002.10826v32020Spatiotemporal Pyramid Network for Video Action Recognition
Yunbo Wang, Mingsheng Long, Jianmin Wang +1
cs.CVarXiv:1903.01038v12019StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
Roberto Henschel, Levon Khachatryan, Hayk Poghosyan +5
cs.CVcs.AIcs.CLarXiv:2403.14773v22024Few-Shot Segmentation Without Meta-Learning: A Good Transductive Inference Is All You Need?
Malik Boudiaf, Hoel Kervadec, Ziko Imtiaz Masud +3
cs.CVarXiv:2012.06166v22020CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Zirui Wang, Mengzhou Xia, Luxi He +10
cs.CLcs.CVarXiv:2406.18521v12024Natural Synthetic Anomalies for Self-Supervised Anomaly Detection and Localization
Hannah M. Schlüter, Jeremy Tan, Benjamin Hou +1
cs.CVarXiv:2109.15222v32021Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
Grace Luo, Lisa Dunlap, Dong Huk Park +2
cs.CVarXiv:2305.14334v22023Targeting Ultimate Accuracy: Face Recognition via Deep Embedding
Jingtuo Liu, Yafeng Deng, Tao Bai +2
cs.CVarXiv:1506.07310v42015Progressive Domain Expansion Network for Single Domain Generalization
Lei Li, Ke Gao, Juan Cao +6
cs.CVarXiv:2103.16050v12021Tracking Everything Everywhere All at Once
Qianqian Wang, Yen-Yu Chang, Ruojin Cai +4
cs.CVarXiv:2306.05422v22023SoccerNet-v2: A Dataset and Benchmarks for Holistic Understanding of Broadcast Soccer Videos
Adrien Deliège, Anthony Cioppa, Silvio Giancola +6
cs.CVarXiv:2011.13367v320203D-LaneNet: End-to-End 3D Multiple Lane Detection
Noa Garnett, Rafi Cohen, Tomer Pe'er +2
cs.CVcs.LGcs.ROarXiv:1811.10203v32018ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Kevin Qinghong Lin, Linjie Li, Difei Gao +6
cs.CVcs.AIcs.CLarXiv:2411.17465v12024DDP: Diffusion Model for Dense Visual Prediction
Yuanfeng Ji, Zhe Chen, Enze Xie +6
cs.CVarXiv:2303.17559v22023Cross-domain Object Detection through Coarse-to-Fine Feature Adaptation
Yangtao Zheng, Di Huang, Songtao Liu +1
cs.CVarXiv:2003.10275v12020Logit Standardization in Knowledge Distillation
Shangquan Sun, Wenqi Ren, Jingzhi Li +2
cs.CVarXiv:2403.01427v12024Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects
Martin Rünz, Lourdes Agapito
cs.CVarXiv:1706.06629v12017R2GenGPT: Radiology Report Generation with Frozen LLMs
Zhanyu Wang, Lingqiao Liu, Lei Wang +1
cs.CVarXiv:2309.09812v22023DDD17: End-To-End DAVIS Driving Dataset
Jonathan Binas, Daniel Neil, Shih-Chii Liu +1
cs.CVarXiv:1711.01458v12017SparseFool: a few pixels make a big difference
Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, Pascal Frossard
cs.CVcs.CRcs.LGarXiv:1811.02248v42018Integrating Triaxial IMU Sensors and Ensemble Learning for Effective Parkinson Disease Severity Classification
Rehan Khan, Muhammad Junaid Asif, Rana Fayyaz Ahmad
cs.AIcs.CVarXiv:2608.28602v12026Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan +1
cs.CVcs.GRarXiv:2311.18482v12023MedTVL: Harnessing Vision and Language for Medical Time Series Classification
Jiexia Ye, Jia Li, Fugee Tsung
cs.AIcs.CVarXiv:2608.28605v12026Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Feng Liu, Shiwei Zhang, Xiaofeng Wang +6
cs.CVarXiv:2411.19108v22024Local Learning Matters: Rethinking Data Heterogeneity in Federated Learning
Matias Mendieta, Taojiannan Yang, Pu Wang +3
cs.LGcs.CVcs.DCarXiv:2111.14213v32021Depth-based hand pose estimation: methods, data, and challenges
James Steven Supancic, Gregory Rogez, Yi Yang +2
cs.CVarXiv:1504.06378v22015The Potential of Haptic Foundation Models
Jianquan Wang, Haiwei Dong, Abdulmotaleb El Saddik
cs.ROcs.CVcs.MMarXiv:2608.28664v12026Generating Holistic 3D Human Motion from Speech
Hongwei Yi, Hualin Liang, Yifei Liu +5
cs.CVcs.GRarXiv:2212.04420v22022Multi-exposure HDR Imaging: A Review of Pixel-level and Feature-level Reconstruction Methods
Qian Tao, Wei Wang, Chaobing Zheng +1
cs.CVarXiv:2608.28674v12026Query-Dependent Video Representation for Moment Retrieval and Highlight Detection
WonJun Moon, Sangeek Hyun, SangUk Park +2
cs.CVcs.AIarXiv:2303.13874v12023Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
Ngan Le, Vidhiwar Singh Rathour, Kashu Yamazaki +2
cs.CVcs.AIarXiv:2108.11510v12021Review of Large Vision Models and Visual Prompt Engineering
Jiaqi Wang, Zhengliang Liu, Lin Zhao +18
cs.CVcs.AIarXiv:2307.00855v12023A Bayesian Data Augmentation Approach for Learning Deep Models
Toan Tran, Trung Pham, Gustavo Carneiro +2
cs.CVcs.LGarXiv:1710.10564v12017Driving Policy Transfer via Modularity and Abstraction
Matthias Müller, Alexey Dosovitskiy, Bernard Ghanem +1
cs.ROcs.CVcs.LGarXiv:1804.09364v32018Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning
Yingdong Hu, Fanqi Lin, Tong Zhang +2
cs.ROcs.AIcs.CLarXiv:2311.17842v22023M$^2$BEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation
Enze Xie, Zhiding Yu, Daquan Zhou +5
cs.CVarXiv:2204.05088v22022PI-RCNN: An Efficient Multi-sensor 3D Object Detector with Point-based Attentive Cont-conv Fusion Module
Liang Xie, Chao Xiang, Zhengxu Yu +4
cs.CVarXiv:1911.06084v32019Referring Image Segmentation via Cross-Modal Progressive Comprehension
Shaofei Huang, Tianrui Hui, Si Liu +5
cs.CVcs.CLarXiv:2010.00514v12020BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects
Bowen Wen, Jonathan Tremblay, Valts Blukis +6
cs.CVcs.AIcs.GRarXiv:2303.14158v12023When Image Denoising Meets High-Level Vision Tasks: A Deep Learning Approach
Ding Liu, Bihan Wen, Xianming Liu +2
cs.CVarXiv:1706.04284v32017Joint Monocular 3D Vehicle Detection and Tracking
Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang +5
cs.CVarXiv:1811.10742v32018Fooling Neural Network Interpretations via Adversarial Model Manipulation
Juyeon Heo, Sunghwan Joo, Taesup Moon
cs.LGcs.AIcs.CVarXiv:1902.02041v32019Actor-Centric Relation Network
Chen Sun, Abhinav Shrivastava, Carl Vondrick +3
cs.CVarXiv:1807.10982v12018Fracture Detection in Pediatric Wrist Trauma X-ray Images Using YOLOv8 Algorithm
Rui-Yang Ju, Weiming Cai
cs.CVarXiv:2304.05071v52023Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
Xintong Wang, Jingheng Pan, Liang Ding +1
cs.CVcs.AIcs.CLarXiv:2403.18715v22024