Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,281 to 8,340 of 18,866

  1. Adapting Segment Anything Model for Change Detection in HR Remote Sensing Images

    Lei Ding, Kun Zhu, Daifeng Peng +3

    cs.CVarXiv:2309.01429v42023
  2. CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation

    Adonay Demewez Gebremedhin, Wessam Shehieb, Sara Alansari +4

    cs.CVarXiv:2608.30758v12026
  3. Unsupervised Domain Adaptation using Generative Adversarial Networks for Semantic Segmentation of Aerial Images

    Bilel Benjdira, Yakoub Bazi, Anis Koubaa +1

    cs.CVarXiv:1905.03198v12019
  4. UFPR-PEs: A Brazilian Face Recognition Benchmark with Self-Declared Race/Color Labels

    Alexandre Diano, Bernardo Biesseck, Gabriel Polo +4

    cs.CVarXiv:2608.30688v12026
  5. PointGrow: Autoregressively Learned Point Cloud Generation with Self-Attention

    Yongbin Sun, Yue Wang, Ziwei Liu +2

    cs.CVarXiv:1810.05591v32018
  6. diffGrad: An Optimization Method for Convolutional Neural Networks

    Shiv Ram Dubey, Soumendu Chakraborty, Swalpa Kumar Roy +3

    cs.LGcs.CVcs.NEarXiv:1909.11015v42019
  7. TUE-Detector: A Tool-Using Expert MLLM-Based Detector for AI-Generated Videos

    Yichen Wu, Haoxuan Qu, Yongxing Dai +5

    cs.CVarXiv:2608.30704v12026
  8. Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation

    Xiaochuan Fan, Kang Zheng, Yuewei Lin +1

    cs.CVarXiv:1504.07159v12015
  9. Prime Sample Attention in Object Detection

    Yuhang Cao, Kai Chen, Chen Change Loy +1

    cs.CVarXiv:1904.04821v22019
  10. Relaxed Transformer Decoders for Direct Action Proposal Generation

    Jing Tan, Jiaqi Tang, Limin Wang +1

    cs.CVarXiv:2102.01894v32021
  11. Swin2SR: SwinV2 Transformer for Compressed Image Super-Resolution and Restoration

    Marcos V. Conde, Ui-Jin Choi, Maxime Burchi +1

    cs.CVeess.IVarXiv:2209.11345v12022
  12. Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations

    Dan Hendrycks, Thomas G. Dietterich

    cs.LGcs.AIcs.CVarXiv:1807.01697v52018
  13. RWF-2000: An Open Large Scale Video Database for Violence Detection

    Ming Cheng, Kunjing Cai, Ming Li

    cs.CVarXiv:1911.05913v32019
  14. AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

    Huawei Wei, Zejun Yang, Zhisheng Wang

    cs.CVcs.GReess.IVarXiv:2403.17694v12024
  15. GasHis-Transformer: A Multi-scale Visual Transformer Approach for Gastric Histopathological Image Detection

    Haoyuan Chen, Chen Li, Ge Wang +9

    cs.CVarXiv:2104.14528v72021
  16. SDM-NET: Deep Generative Network for Structured Deformable Mesh

    Lin Gao, Jie Yang, Tong Wu +4

    cs.GRcs.CVarXiv:1908.04520v22019
  17. SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

    Ranjit Raut, Aarav Subedi, Sagun Rai +1

    cs.AIcs.CVarXiv:2609.00018v12026
  18. Fake it till you make it: Learning transferable representations from synthetic ImageNet clones

    Mert Bulent Sariyildiz, Karteek Alahari, Diane Larlus +1

    cs.CVcs.LGarXiv:2212.08420v22022
  19. RelTR: Relation Transformer for Scene Graph Generation

    Yuren Cong, Michael Ying Yang, Bodo Rosenhahn

    cs.CVarXiv:2201.11460v32022
  20. Wide-Slice Residual Networks for Food Recognition

    Niki Martinel, Gian Luca Foresti, Christian Micheloni

    cs.CVarXiv:1612.06543v12016
  21. ObjectFormer for Image Manipulation Detection and Localization

    Junke Wang, Zuxuan Wu, Jingjing Chen +4

    cs.CVarXiv:2203.14681v22022
  22. Humble Teachers Teach Better Students for Semi-Supervised Object Detection

    Yihe Tang, Weifeng Chen, Yijun Luo +1

    cs.CVarXiv:2106.10456v12021
  23. Blind Face Restoration via Deep Multi-scale Component Dictionaries

    Xiaoming Li, Chaofeng Chen, Shangchen Zhou +3

    cs.CVarXiv:2008.00418v12020
  24. Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

    Zhe Gan, Linjie Li, Chunyuan Li +3

    cs.CVcs.CLarXiv:2210.09263v12022
  25. Task Driven Generative Modeling for Unsupervised Domain Adaptation: Application to X-ray Image Segmentation

    Yue Zhang, Shun Miao, Tommaso Mansi +1

    cs.CVarXiv:1806.07201v12018
  26. Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction

    Hansheng Chen, Jiatao Gu, Anpei Chen +4

    cs.CVarXiv:2304.06714v42023
  27. Cross-View Image Matching for Geo-localization in Urban Environments

    Yicong Tian, Chen Chen, Mubarak Shah

    cs.CVarXiv:1703.07815v12017
  28. Complexer-YOLO: Real-Time 3D Object Detection and Tracking on Semantic Point Clouds

    Martin Simon, Karl Amende, Andrea Kraus +5

    cs.CVarXiv:1904.07537v12019
  29. Learning Policies for Adaptive Tracking with Deep Feature Cascades

    Chen Huang, Simon Lucey, Deva Ramanan

    cs.CVarXiv:1708.02973v22017
  30. Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training

    Shenggui Li, Hongxin Liu, Zhengda Bian +5

    cs.LGcs.AIcs.CLarXiv:2110.14883v32021
  31. Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis

    Konstantinos Moutselos, Ilias Maglogiannis

    cs.CVcs.AIarXiv:2608.30835v12026
  32. What Makes Good Synthetic Training Data for Learning Disparity and Optical Flow Estimation?

    Nikolaus Mayer, Eddy Ilg, Philipp Fischer +4

    cs.CVstat.MLarXiv:1801.06397v32018
  33. Disentangle Your Dense Object Detector

    Zehui Chen, Chenhongyi Yang, Qiaofei Li +3

    cs.CVarXiv:2107.02963v22021
  34. TPNet: Trajectory Proposal Network for Motion Prediction

    Liangji Fang, Qinhong Jiang, Jianping Shi +1

    cs.CVarXiv:2004.12255v22020
  35. Pruning Neural Networks at Initialization: Why are We Missing the Mark?

    Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy +1

    cs.LGcs.CVcs.NEarXiv:2009.08576v22020
  36. Non-Stationary Texture Synthesis by Adversarial Expansion

    Yang Zhou, Zhen Zhu, Xiang Bai +3

    cs.GRcs.CVarXiv:1805.04487v12018
  37. A graph-transformer for whole slide image classification

    Yi Zheng, Rushin H. Gindra, Emily J. Green +4

    cs.CVarXiv:2205.09671v12022
  38. ReenactGAN: Learning to Reenact Faces via Boundary Transfer

    Wayne Wu, Yunxuan Zhang, Cheng Li +2

    cs.CVcs.AIcs.GRarXiv:1807.11079v12018
  39. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration

    Junyang Wang, Haiyang Xu, Haitao Jia +6

    cs.CLcs.CVarXiv:2406.01014v12024
  40. SMPLicit: Topology-aware Generative Model for Clothed People

    Enric Corona, Albert Pumarola, Guillem Alenyà +2

    cs.CVarXiv:2103.06871v22021
  41. DSVT: Dynamic Sparse Voxel Transformer with Rotated Sets

    Haiyang Wang, Chen Shi, Shaoshuai Shi +5

    cs.CVarXiv:2301.06051v22023
  42. Unsupervised Learning of Visual Representations using Videos

    Xiaolong Wang, Abhinav Gupta

    cs.CVarXiv:1505.00687v22015
  43. Identity-Conditioned Latent Consistency Distillation for Face Synthesis

    Tiago Kienen Chaves, Bernardo Biesseck, David Menotti

    cs.CVarXiv:2608.31053v12026
  44. Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers

    Lei Ke, Yu-Wing Tai, Chi-Keung Tang

    cs.CVarXiv:2103.12340v12021
  45. PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model

    Juncai Peng, Yi Liu, Shiyu Tang +13

    cs.CVcs.AIarXiv:2204.02681v12022
  46. Fast Feature Fool: A data independent approach to universal adversarial perturbations

    Konda Reddy Mopuri, Utsav Garg, R. Venkatesh Babu

    cs.CVarXiv:1707.05572v12017
  47. General Framework to Evaluate Unlinkability in Biometric Template Protection Systems

    Marta Gomez-Barrero, Javier Galbally, Christian Rathgeb +1

    cs.CVarXiv:2311.04633v12023
  48. Specificity-preserving RGB-D Saliency Detection

    Tao Zhou, Deng-Ping Fan, Geng Chen +2

    cs.CVarXiv:2108.08162v22021
  49. TallyQA: Answering Complex Counting Questions

    Manoj Acharya, Kushal Kafle, Christopher Kanan

    cs.CVarXiv:1810.12440v22018
  50. Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction

    Osama Makansi, Eddy Ilg, Özgün Cicek +1

    cs.CVarXiv:1906.03631v22019
  51. Multi-Task Recurrent Convolutional Network with Correlation Loss for Surgical Video Analysis

    Yueming Jin, Huaxia Li, Qi Dou +4

    cs.CVcs.LGeess.IVarXiv:1907.06099v12019
  52. Distort-and-Recover: Color Enhancement using Deep Reinforcement Learning

    Jongchan Park, Joon-Young Lee, Donggeun Yoo +1

    cs.CVarXiv:1804.04450v22018
  53. RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation

    Quan Hao, Ziyang Tao, Chenxi Zhang +3

    cs.CVcs.AIarXiv:2608.30727v12026
  54. Spatiotemporal Inconsistency Learning for DeepFake Video Detection

    Zhihao Gu, Yang Chen, Taiping Yao +4

    cs.CVarXiv:2109.01860v32021
  55. CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

    Yuqi Lin, Minghao Chen, Wenxiao Wang +5

    cs.CVcs.AIarXiv:2212.09506v32022
  56. Deep convolutional neural networks for predominant instrument recognition in polyphonic music

    Yoonchang Han, Jaehun Kim, Kyogu Lee

    cs.SDcs.CVcs.LGarXiv:1605.09507v32016
  57. Unsupervised Image Captioning

    Yang Feng, Lin Ma, Wei Liu +1

    cs.CVarXiv:1811.10787v22018
  58. Cost-efficient Active Learning for Referring Image Segmentation and Grounding

    Junbeom Hong, Seonghoon Yu, Hyung Rok Jung +2

    cs.CVcs.AIarXiv:2608.30621v22026
  59. Divide and Grow: Capturing Huge Diversity in Crowd Images with Incrementally Growing CNN

    Deepak Babu Sam, Neeraj N Sajjan, R. Venkatesh Babu

    cs.CVarXiv:1807.09993v12018
  60. InstanceDiffusion: Instance-level Control for Image Generation

    Xudong Wang, Trevor Darrell, Sai Saketh Rambhatla +2

    cs.CVcs.AIcs.LGarXiv:2402.03290v12024