Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,281 to 8,340 of 18,866
Adapting Segment Anything Model for Change Detection in HR Remote Sensing Images
Lei Ding, Kun Zhu, Daifeng Peng +3
cs.CVarXiv:2309.01429v42023CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
Adonay Demewez Gebremedhin, Wessam Shehieb, Sara Alansari +4
cs.CVarXiv:2608.30758v12026Unsupervised Domain Adaptation using Generative Adversarial Networks for Semantic Segmentation of Aerial Images
Bilel Benjdira, Yakoub Bazi, Anis Koubaa +1
cs.CVarXiv:1905.03198v12019UFPR-PEs: A Brazilian Face Recognition Benchmark with Self-Declared Race/Color Labels
Alexandre Diano, Bernardo Biesseck, Gabriel Polo +4
cs.CVarXiv:2608.30688v12026PointGrow: Autoregressively Learned Point Cloud Generation with Self-Attention
Yongbin Sun, Yue Wang, Ziwei Liu +2
cs.CVarXiv:1810.05591v32018diffGrad: An Optimization Method for Convolutional Neural Networks
Shiv Ram Dubey, Soumendu Chakraborty, Swalpa Kumar Roy +3
cs.LGcs.CVcs.NEarXiv:1909.11015v42019TUE-Detector: A Tool-Using Expert MLLM-Based Detector for AI-Generated Videos
Yichen Wu, Haoxuan Qu, Yongxing Dai +5
cs.CVarXiv:2608.30704v12026Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation
Xiaochuan Fan, Kang Zheng, Yuewei Lin +1
cs.CVarXiv:1504.07159v12015Prime Sample Attention in Object Detection
Yuhang Cao, Kai Chen, Chen Change Loy +1
cs.CVarXiv:1904.04821v22019Relaxed Transformer Decoders for Direct Action Proposal Generation
Jing Tan, Jiaqi Tang, Limin Wang +1
cs.CVarXiv:2102.01894v32021Swin2SR: SwinV2 Transformer for Compressed Image Super-Resolution and Restoration
Marcos V. Conde, Ui-Jin Choi, Maxime Burchi +1
cs.CVeess.IVarXiv:2209.11345v12022Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
Dan Hendrycks, Thomas G. Dietterich
cs.LGcs.AIcs.CVarXiv:1807.01697v52018RWF-2000: An Open Large Scale Video Database for Violence Detection
Ming Cheng, Kunjing Cai, Ming Li
cs.CVarXiv:1911.05913v32019AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Huawei Wei, Zejun Yang, Zhisheng Wang
cs.CVcs.GReess.IVarXiv:2403.17694v12024GasHis-Transformer: A Multi-scale Visual Transformer Approach for Gastric Histopathological Image Detection
Haoyuan Chen, Chen Li, Ge Wang +9
cs.CVarXiv:2104.14528v72021SDM-NET: Deep Generative Network for Structured Deformable Mesh
Lin Gao, Jie Yang, Tong Wu +4
cs.GRcs.CVarXiv:1908.04520v22019SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces
Ranjit Raut, Aarav Subedi, Sagun Rai +1
cs.AIcs.CVarXiv:2609.00018v12026Fake it till you make it: Learning transferable representations from synthetic ImageNet clones
Mert Bulent Sariyildiz, Karteek Alahari, Diane Larlus +1
cs.CVcs.LGarXiv:2212.08420v22022RelTR: Relation Transformer for Scene Graph Generation
Yuren Cong, Michael Ying Yang, Bodo Rosenhahn
cs.CVarXiv:2201.11460v32022Wide-Slice Residual Networks for Food Recognition
Niki Martinel, Gian Luca Foresti, Christian Micheloni
cs.CVarXiv:1612.06543v12016ObjectFormer for Image Manipulation Detection and Localization
Junke Wang, Zuxuan Wu, Jingjing Chen +4
cs.CVarXiv:2203.14681v22022Humble Teachers Teach Better Students for Semi-Supervised Object Detection
Yihe Tang, Weifeng Chen, Yijun Luo +1
cs.CVarXiv:2106.10456v12021Blind Face Restoration via Deep Multi-scale Component Dictionaries
Xiaoming Li, Chaofeng Chen, Shangchen Zhou +3
cs.CVarXiv:2008.00418v12020Vision-Language Pre-training: Basics, Recent Advances, and Future Trends
Zhe Gan, Linjie Li, Chunyuan Li +3
cs.CVcs.CLarXiv:2210.09263v12022Task Driven Generative Modeling for Unsupervised Domain Adaptation: Application to X-ray Image Segmentation
Yue Zhang, Shun Miao, Tommaso Mansi +1
cs.CVarXiv:1806.07201v12018Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction
Hansheng Chen, Jiatao Gu, Anpei Chen +4
cs.CVarXiv:2304.06714v42023Cross-View Image Matching for Geo-localization in Urban Environments
Yicong Tian, Chen Chen, Mubarak Shah
cs.CVarXiv:1703.07815v12017Complexer-YOLO: Real-Time 3D Object Detection and Tracking on Semantic Point Clouds
Martin Simon, Karl Amende, Andrea Kraus +5
cs.CVarXiv:1904.07537v12019Learning Policies for Adaptive Tracking with Deep Feature Cascades
Chen Huang, Simon Lucey, Deva Ramanan
cs.CVarXiv:1708.02973v22017Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training
Shenggui Li, Hongxin Liu, Zhengda Bian +5
cs.LGcs.AIcs.CLarXiv:2110.14883v32021Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis
Konstantinos Moutselos, Ilias Maglogiannis
cs.CVcs.AIarXiv:2608.30835v12026What Makes Good Synthetic Training Data for Learning Disparity and Optical Flow Estimation?
Nikolaus Mayer, Eddy Ilg, Philipp Fischer +4
cs.CVstat.MLarXiv:1801.06397v32018Disentangle Your Dense Object Detector
Zehui Chen, Chenhongyi Yang, Qiaofei Li +3
cs.CVarXiv:2107.02963v22021TPNet: Trajectory Proposal Network for Motion Prediction
Liangji Fang, Qinhong Jiang, Jianping Shi +1
cs.CVarXiv:2004.12255v22020Pruning Neural Networks at Initialization: Why are We Missing the Mark?
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy +1
cs.LGcs.CVcs.NEarXiv:2009.08576v22020Non-Stationary Texture Synthesis by Adversarial Expansion
Yang Zhou, Zhen Zhu, Xiang Bai +3
cs.GRcs.CVarXiv:1805.04487v12018A graph-transformer for whole slide image classification
Yi Zheng, Rushin H. Gindra, Emily J. Green +4
cs.CVarXiv:2205.09671v12022ReenactGAN: Learning to Reenact Faces via Boundary Transfer
Wayne Wu, Yunxuan Zhang, Cheng Li +2
cs.CVcs.AIcs.GRarXiv:1807.11079v12018Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
Junyang Wang, Haiyang Xu, Haitao Jia +6
cs.CLcs.CVarXiv:2406.01014v12024SMPLicit: Topology-aware Generative Model for Clothed People
Enric Corona, Albert Pumarola, Guillem Alenyà +2
cs.CVarXiv:2103.06871v22021DSVT: Dynamic Sparse Voxel Transformer with Rotated Sets
Haiyang Wang, Chen Shi, Shaoshuai Shi +5
cs.CVarXiv:2301.06051v22023Unsupervised Learning of Visual Representations using Videos
Xiaolong Wang, Abhinav Gupta
cs.CVarXiv:1505.00687v22015Identity-Conditioned Latent Consistency Distillation for Face Synthesis
Tiago Kienen Chaves, Bernardo Biesseck, David Menotti
cs.CVarXiv:2608.31053v12026Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
cs.CVarXiv:2103.12340v12021PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model
Juncai Peng, Yi Liu, Shiyu Tang +13
cs.CVcs.AIarXiv:2204.02681v12022Fast Feature Fool: A data independent approach to universal adversarial perturbations
Konda Reddy Mopuri, Utsav Garg, R. Venkatesh Babu
cs.CVarXiv:1707.05572v12017General Framework to Evaluate Unlinkability in Biometric Template Protection Systems
Marta Gomez-Barrero, Javier Galbally, Christian Rathgeb +1
cs.CVarXiv:2311.04633v12023Specificity-preserving RGB-D Saliency Detection
Tao Zhou, Deng-Ping Fan, Geng Chen +2
cs.CVarXiv:2108.08162v22021TallyQA: Answering Complex Counting Questions
Manoj Acharya, Kushal Kafle, Christopher Kanan
cs.CVarXiv:1810.12440v22018Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
Osama Makansi, Eddy Ilg, Özgün Cicek +1
cs.CVarXiv:1906.03631v22019Multi-Task Recurrent Convolutional Network with Correlation Loss for Surgical Video Analysis
Yueming Jin, Huaxia Li, Qi Dou +4
cs.CVcs.LGeess.IVarXiv:1907.06099v12019Distort-and-Recover: Color Enhancement using Deep Reinforcement Learning
Jongchan Park, Joon-Young Lee, Donggeun Yoo +1
cs.CVarXiv:1804.04450v22018RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation
Quan Hao, Ziyang Tao, Chenxi Zhang +3
cs.CVcs.AIarXiv:2608.30727v12026Spatiotemporal Inconsistency Learning for DeepFake Video Detection
Zhihao Gu, Yang Chen, Taiping Yao +4
cs.CVarXiv:2109.01860v32021CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation
Yuqi Lin, Minghao Chen, Wenxiao Wang +5
cs.CVcs.AIarXiv:2212.09506v32022Deep convolutional neural networks for predominant instrument recognition in polyphonic music
Yoonchang Han, Jaehun Kim, Kyogu Lee
cs.SDcs.CVcs.LGarXiv:1605.09507v32016Unsupervised Image Captioning
Yang Feng, Lin Ma, Wei Liu +1
cs.CVarXiv:1811.10787v22018Cost-efficient Active Learning for Referring Image Segmentation and Grounding
Junbeom Hong, Seonghoon Yu, Hyung Rok Jung +2
cs.CVcs.AIarXiv:2608.30621v22026Divide and Grow: Capturing Huge Diversity in Crowd Images with Incrementally Growing CNN
Deepak Babu Sam, Neeraj N Sajjan, R. Venkatesh Babu
cs.CVarXiv:1807.09993v12018InstanceDiffusion: Instance-level Control for Image Generation
Xudong Wang, Trevor Darrell, Sai Saketh Rambhatla +2
cs.CVcs.AIcs.LGarXiv:2402.03290v12024