Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,981 to 2,040 of 18,837

  1. Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras

    Lingni Ma, Jörg Stückler, Christian Kerl +1

    cs.CVarXiv:1703.08866v22017
  2. Deep Learning in Photoacoustic Tomography: Current approaches and future directions

    Andreas Hauptmann, Ben Cox

    eess.IVcs.CVcs.LGarXiv:2009.07608v12020
  3. The Elements of End-to-end Deep Face Recognition: A Survey of Recent Advances

    Hang Du, Hailin Shi, Dan Zeng +2

    cs.CVarXiv:2009.13290v42020
  4. SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

    Yuncong Yang, Zhengtao Han, Furkan Ozyurt +6

    cs.CVarXiv:2609.09155v12026
  5. MDFN: Multi-Scale Deep Feature Learning Network for Object Detection

    Wenchi Ma, Yuanwei Wu, Feng Cen +1

    cs.CVarXiv:1912.04514v12019
  6. Myocardial Strain Drift Correction in Deep Learning Based Ultrasound Tracking

    Thierry Judge, Nicolas Duchateau, Andreas Østvik +6

    eess.IVcs.AIcs.CVarXiv:2609.09577v12026
  7. NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates Space

    Jiawei Yao, Chuming Li, Keqiang Sun +4

    cs.CVarXiv:2309.14616v32023
  8. DA-GAN: Instance-level Image Translation by Deep Attention Generative Adversarial Networks (with Supplementary Materials)

    Shuang Ma, Jianlong Fu, Chang Wen Chen +1

    cs.CVarXiv:1802.06454v12018
  9. DuAT: Dual-Aggregation Transformer Network for Medical Image Segmentation

    Feilong Tang, Qiming Huang, Jinfeng Wang +3

    cs.CVarXiv:2212.11677v12022
  10. AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

    Zehui Chen, Zhenyu Li, Shiquan Zhang +5

    cs.CVarXiv:2201.06493v22022
  11. GANVO: Unsupervised Deep Monocular Visual Odometry and Depth Estimation with Generative Adversarial Networks

    Yasin Almalioglu, Muhamad Risqi U. Saputra, Pedro P. B. de Gusmao +2

    cs.LGcs.CVstat.MLarXiv:1809.05786v32018
  12. Visual Causal Feature Learning

    Krzysztof Chalupka, Pietro Perona, Frederick Eberhardt

    stat.MLcs.AIcs.CVarXiv:1412.2309v22014
  13. Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain Adaptation

    Guo-Sen Xie, Xu-Yao Zhang, Shuicheng Yan +1

    cs.CVarXiv:1601.07977v12016
  14. Methods of Hierarchical Clustering

    Fionn Murtagh, Pedro Contreras

    cs.IRcs.CVmath.STarXiv:1105.0121v12011
  15. Fully Convolutional Networks for Diabetic Foot Ulcer Segmentation

    Manu Goyal, Neil D. Reeves, Satyan Rajbhandari +2

    cs.CVarXiv:1708.01928v12017
  16. Edge-Host Partitioning of Deep Neural Networks with Feature Space Encoding for Resource-Constrained Internet-of-Things Platforms

    Jong Hwan Ko, Taesik Na, Mohammad Faisal Amir +1

    cs.CVarXiv:1802.03835v12018
  17. Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?

    Rishi Bommasani, Kathleen A. Creel, Ananya Kumar +2

    cs.LGcs.AIcs.CLarXiv:2211.13972v12022
  18. How Deep Learning Sees the World: A Survey on Adversarial Attacks & Defenses

    Joana C. Costa, Tiago Roxo, Hugo Proença +1

    cs.CVarXiv:2305.10862v12023
  19. Image Forgery Localization Based on Multi-Scale Convolutional Neural Networks

    Yaqi Liu, Qingxiao Guan, Xianfeng Zhao +1

    cs.CVcs.MMarXiv:1706.07842v42017
  20. PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report

    Andrey Ignatov, Radu Timofte, Thang Van Vu +45

    cs.CVarXiv:1810.01641v12018
  21. OmniTact: A Multi-Directional High Resolution Touch Sensor

    Akhil Padmanabha, Frederik Ebert, Stephen Tian +3

    cs.ROcs.CVcs.LGarXiv:2003.06965v12020
  22. Spectral Superresolution of Multispectral Imagery with Joint Sparse and Low-Rank Learning

    Lianru Gao, Danfeng Hong, Jing Yao +3

    eess.IVcs.CVarXiv:2007.14006v12020
  23. Long Movie Clip Classification with State-Space Video Models

    Md Mohaiminul Islam, Gedas Bertasius

    cs.CVarXiv:2204.01692v32022
  24. Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models

    Wen Wang, Yan Jiang, Kangyang Xie +5

    cs.CVarXiv:2303.17599v32023
  25. VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

    Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6

    cs.CVcs.AIarXiv:2609.09396v12026
  26. Self-supervised Moving Vehicle Tracking with Stereo Sound

    Chuang Gan, Hang Zhao, Peihao Chen +2

    cs.CVcs.LGcs.SDarXiv:1910.11760v12019
  27. Can stable and accurate neural networks be computed? -- On the barriers of deep learning and Smale's 18th problem

    Matthew J. Colbrook, Vegard Antun, Anders C. Hansen

    cs.LGcs.CVcs.NEarXiv:2101.08286v22021
  28. Real-Time Drone Detection and Tracking With Visible, Thermal and Acoustic Sensors

    Fredrik Svanstrom, Cristofer Englund, Fernando Alonso-Fernandez

    cs.CVeess.SParXiv:2007.07396v22020
  29. Learning with Privileged Information for Efficient Image Super-Resolution

    Wonkyung Lee, Junghyup Lee, Dohyung Kim +1

    cs.CVarXiv:2007.07524v12020
  30. No Free Checker: A Survey of Verifiers for Robot Policies

    Yang Wan, Xihang Yue, Zhirui Liu +7

    cs.ROcs.AIcs.CVarXiv:2609.09250v12026
  31. Affective EEG-Based Person Identification Using the Deep Learning Approach

    Theerawit Wilaiprasitporn, Apiwat Ditthapron, Karis Matchaparn +3

    eess.SPcs.CVarXiv:1807.03147v32018
  32. Relightable Gaussian Codec Avatars

    Shunsuke Saito, Gabriel Schwartz, Tomas Simon +2

    cs.GRcs.CVarXiv:2312.03704v22023
  33. A General Decoupled Learning Framework for Parameterized Image Operators

    Qingnan Fan, Dongdong Chen, Lu Yuan +3

    cs.CVarXiv:1907.05852v12019
  34. The Limitations of Adversarial Training and the Blind-Spot Attack

    Huan Zhang, Hongge Chen, Zhao Song +3

    stat.MLcs.CRcs.CVarXiv:1901.04684v12019
  35. Low-Rank Pairwise Alignment Bilinear Network For Few-Shot Fine-Grained Image Classification

    Huaxi Huang, Junjie Zhang, Jian Zhang +2

    cs.CVarXiv:1908.01313v32019
  36. Diffusion Probabilistic Model Made Slim

    Xingyi Yang, Daquan Zhou, Jiashi Feng +1

    cs.CVeess.IVarXiv:2211.17106v12022
  37. JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

    Yiyang Ma, Xingchao Liu, Xiaokang Chen +11

    cs.CVcs.AIcs.CLarXiv:2411.07975v22024
  38. Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

    Xiangming Gu, Xiaosen Zheng, Tianyu Pang +5

    cs.CLcs.CRcs.CVarXiv:2402.08567v22024
  39. On The Convergence of Gradient Descent for Finding the Riemannian Center of Mass

    Bijan Afsari, Roberto Tron, René Vidal

    math.DGcs.CVmath.NAarXiv:1201.0925v12011
  40. How to Read Paintings: Semantic Art Understanding with Multi-Modal Retrieval

    Noa Garcia, George Vogiatzis

    cs.CVarXiv:1810.09617v12018
  41. LF-YOLO: A Lighter and Faster YOLO for Weld Defect Detection of X-ray Image

    Moyun Liu, Youping Chen, Lei He +2

    cs.CVarXiv:2110.15045v22021
  42. IMos: Intent-Driven Full-Body Motion Synthesis for Human-Object Interactions

    Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik +2

    cs.CVcs.GRcs.LGarXiv:2212.07555v32022
  43. Deep feature compression for collaborative object detection

    Hyomin Choi, Ivan V. Bajic

    cs.CVarXiv:1802.03931v12018
  44. Jailbreaking Attack against Multimodal Large Language Model

    Zhenxing Niu, Haodong Ren, Xinbo Gao +2

    cs.LGcs.CLcs.CRarXiv:2402.02309v12024
  45. DiffuseVAE: Efficient, Controllable and High-Fidelity Generation from Low-Dimensional Latents

    Kushagra Pandey, Avideep Mukherjee, Piyush Rai +1

    cs.LGcs.CVarXiv:2201.00308v32022
  46. Multi-Layer Pseudo-Supervision for Histopathology Tissue Semantic Segmentation using Patch-level Classification Labels

    Chu Han, Jiatai Lin, Jinhai Mai +15

    eess.IVcs.CVq-bio.QMarXiv:2110.08048v12021
  47. Adversarial Examples that Fool Detectors

    Jiajun Lu, Hussein Sibai, Evan Fabry

    cs.CVcs.AIcs.GRarXiv:1712.02494v12017
  48. Correlation Tracking via Joint Discrimination and Reliability Learning

    Chong Sun, Dong Wang, Huchuan Lu +1

    cs.CVarXiv:1804.08965v12018
  49. PaLI-3 Vision Language Models: Smaller, Faster, Stronger

    Xi Chen, Xiao Wang, Lucas Beyer +16

    cs.CVarXiv:2310.09199v22023
  50. Margin Sample Mining Loss: A Deep Learning Based Method for Person Re-identification

    Qiqi Xiao, Hao Luo, Chi Zhang

    cs.CVarXiv:1710.00478v32017
  51. DE-GAN: A Conditional Generative Adversarial Network for Document Enhancement

    Mohamed Ali Souibgui, Yousri Kessentini

    cs.CVarXiv:2010.08764v12020
  52. Low-Rank Few-Shot Adaptation of Vision-Language Models

    Maxime Zanella, Ismail Ben Ayed

    cs.CVarXiv:2405.18541v22024
  53. Micro-Attention for Micro-Expression recognition

    Chongyang Wang, Min Peng, Tao Bi +1

    cs.CVarXiv:1811.02360v52018
  54. RECALL: Replay-based Continual Learning in Semantic Segmentation

    Andrea Maracani, Umberto Michieli, Marco Toldo +1

    cs.CVarXiv:2108.03673v22021
  55. Generate, Segment and Refine: Towards Generic Manipulation Segmentation

    Peng Zhou, Bor-Chun Chen, Xintong Han +4

    cs.CVarXiv:1811.09729v32018
  56. Creativity: Generating Diverse Questions using Variational Autoencoders

    Unnat Jain, Ziyu Zhang, Alexander Schwing

    cs.CVarXiv:1704.03493v12017
  57. Self6D: Self-Supervised Monocular 6D Object Pose Estimation

    Gu Wang, Fabian Manhardt, Jianzhun Shao +3

    cs.CVarXiv:2004.06468v32020
  58. Monocular Quasi-Dense 3D Object Tracking

    Hou-Ning Hu, Yung-Hsu Yang, Tobias Fischer +3

    cs.CVarXiv:2103.07351v12021
  59. SpaceNet 6: Multi-Sensor All Weather Mapping Dataset

    Jacob Shermeyer, Daniel Hogan, Jason Brown +8

    eess.IVcs.CVarXiv:2004.06500v12020
  60. Combined tract segmentation and orientation mapping for bundle-specific tractography

    Jakob Wasserthal, Peter Neher, Dusan Hirjak +1

    cs.CVarXiv:1901.10271v22019