Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,841 to 12,900 of 18,822

  1. Describing Multimedia Content using Attention-based Encoder--Decoder Networks

    Kyunghyun Cho, Aaron Courville, Yoshua Bengio

    cs.NEcs.CLcs.CVarXiv:1507.01053v12015
  2. KISS-GS: 3D Gaussian Splatting Compression Kept Simple

    Wieland Morgenstern, Friedrich Elias Branschke, Florian Fleischmann +5

    cs.CVarXiv:2608.26948v12026
  3. Behavior Transformers: Cloning $k$ modes with one stone

    Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya +1

    cs.LGcs.AIcs.CVarXiv:2206.11251v22022
  4. Understanding Deep Learning Techniques for Image Segmentation

    Swarnendu Ghosh, Nibaran Das, Ishita Das +1

    cs.CVcs.LGcs.NEarXiv:1907.06119v12019
  5. xView: Objects in Context in Overhead Imagery

    Darius Lam, Richard Kuzma, Kevin McGee +5

    cs.CVarXiv:1802.07856v12018
  6. Differentiable Jitter Correction using Deep Learning-based Image Quality Metric for Phase-Contrast Micro-CT

    Junan Chen, Yiting Jia, Joscha Maier +6

    cs.CVarXiv:2608.27034v12026
  7. Extended Object Tracking: Introduction, Overview and Applications

    Karl Granstrom, Marcus Baum, Stephan Reuter

    cs.CVeess.SPeess.SYarXiv:1604.00970v32016
  8. Recognize Anything: A Strong Image Tagging Model

    Youcai Zhang, Xinyu Huang, Jinyu Ma +9

    cs.CVarXiv:2306.03514v32023
  9. Real-time Action Recognition with Enhanced Motion Vector CNNs

    Bowen Zhang, Limin Wang, Zhe Wang +2

    cs.CVarXiv:1604.07669v12016
  10. Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video

    Xiaowei Zhou, Menglong Zhu, Spyridon Leonardos +2

    cs.CVarXiv:1511.09439v22015
  11. Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval

    Xiu-Shen Wei, Jian-Hao Luo, Jianxin Wu +1

    cs.CVarXiv:1604.04994v22016
  12. Advances in Medical Image Analysis with Vision Transformers: A Comprehensive Review

    Reza Azad, Amirhossein Kazerouni, Moein Heidari +6

    cs.CVarXiv:2301.03505v32023
  13. Combo Loss: Handling Input and Output Imbalance in Multi-Organ Segmentation

    Saeid Asgari Taghanaki, Yefeng Zheng, S. Kevin Zhou +5

    cs.CVarXiv:1805.02798v62018
  14. Virtual iEEG from Scalp EEG: Charting the Landscape of Source Imaging, Intracranial Inference and Reconstruction

    Dongyi He, Xiangkai Wang, Hongjie Yan +3

    cs.CVarXiv:2608.26998v12026
  15. UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation

    Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed +3

    cs.CVarXiv:2212.04497v32022
  16. Some Improvements on Deep Convolutional Neural Network Based Image Classification

    Andrew G. Howard

    cs.CVarXiv:1312.5402v12013
  17. Context-aware Synthesis for Video Frame Interpolation

    Simon Niklaus, Feng Liu

    cs.CVarXiv:1803.10967v12018
  18. Unified Concept Editing in Diffusion Models

    Rohit Gandikota, Hadas Orgad, Yonatan Belinkov +2

    cs.CVcs.LGarXiv:2308.14761v22023
  19. Vector Neurons: A General Framework for SO(3)-Equivariant Networks

    Congyue Deng, Or Litany, Yueqi Duan +3

    cs.CVarXiv:2104.12229v12021
  20. MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

    Junyi Zhang, Charles Herrmann, Junhwa Hur +5

    cs.CVarXiv:2410.03825v22024
  21. End-to-End Variational Networks for Accelerated MRI Reconstruction

    Anuroop Sriram, Jure Zbontar, Tullie Murrell +5

    eess.IVcs.CVarXiv:2004.06688v22020
  22. Multi-Scale Continuous CRFs as Sequential Deep Networks for Monocular Depth Estimation

    Dan Xu, Elisa Ricci, Wanli Ouyang +2

    cs.CVarXiv:1704.02157v12017
  23. FOSTER: Feature Boosting and Compression for Class-Incremental Learning

    Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye +1

    cs.CVcs.LGarXiv:2204.04662v22022
  24. DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion

    Zixiang Zhao, Haowen Bai, Yuanzhi Zhu +7

    cs.CVarXiv:2303.06840v22023
  25. Fruit recognition from images using deep learning

    Horea Mureşan, Mihai Oltean

    cs.CVarXiv:1712.00580v102017
  26. Learning Distilled Collaboration Graph for Multi-Agent Perception

    Yiming Li, Shunli Ren, Pengxiang Wu +3

    cs.CVcs.ROarXiv:2111.00643v22021
  27. SSH: Single Stage Headless Face Detector

    Mahyar Najibi, Pouya Samangouei, Rama Chellappa +1

    cs.CVarXiv:1708.03979v32017
  28. Weakly-Supervised Convolutional Neural Networks for Multimodal Image Registration

    Yipeng Hu, Marc Modat, Eli Gibson +11

    cs.CVcs.AIcs.LGarXiv:1807.03361v12018
  29. Deep-Plant: Plant Identification with convolutional neural networks

    Sue Han Lee, Chee Seng Chan, Paul Wilkin +1

    cs.CVcs.AIcs.NEarXiv:1506.08425v12015
  30. Deep representation learning for human motion prediction and classification

    Judith Bütepage, Michael Black, Danica Kragic +1

    cs.CVarXiv:1702.07486v22017
  31. Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model

    Jiahao Li, Hao Tan, Kai Zhang +7

    cs.CVarXiv:2311.06214v22023
  32. Understanding the Limitations of CNN-based Absolute Camera Pose Regression

    Torsten Sattler, Qunjie Zhou, Marc Pollefeys +1

    cs.CVarXiv:1903.07504v12019
  33. LRRNet: A Novel Representation Learning Guided Fusion Network for Infrared and Visible Images

    Hui Li, Tianyang Xu, Xiao-Jun Wu +2

    cs.CVarXiv:2304.05172v22023
  34. Semi-parametric Topological Memory for Navigation

    Nikolay Savinov, Alexey Dosovitskiy, Vladlen Koltun

    cs.LGcs.AIcs.CVarXiv:1803.00653v12018
  35. WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

    Krishna Srinivasan, Karthik Raman, Jiecao Chen +2

    cs.CVcs.CLcs.IRarXiv:2103.01913v22021
  36. Beyond One-hot Encoding: lower dimensional target embedding

    Pau Rodríguez, Miguel A. Bautista, Jordi Gonzàlez +1

    cs.CVcs.AIarXiv:1806.10805v12018
  37. No New-Net

    Fabian Isensee, Philipp Kickingereder, Wolfgang Wick +2

    cs.CVarXiv:1809.10483v22018
  38. NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

    Gengze Zhou, Yicong Hong, Qi Wu

    cs.CVcs.AIcs.CLarXiv:2305.16986v32023
  39. Video Pixel Networks

    Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan +4

    cs.CVcs.LGarXiv:1610.00527v12016
  40. Video Processing from Electro-optical Sensors for Object Detection and Tracking in Maritime Environment: A Survey

    D. K. Prasad, D. Rajan, L. Rachmawati +2

    cs.CVarXiv:1611.05842v12016
  41. Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

    Agniv Chatterjee, Georgios Pavlakos

    cs.CVarXiv:2608.27407v12026
  42. BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

    Zhiliang Peng, Li Dong, Hangbo Bao +2

    cs.CVarXiv:2208.06366v22022
  43. Population Based Augmentation: Efficient Learning of Augmentation Policy Schedules

    Daniel Ho, Eric Liang, Ion Stoica +2

    cs.CVcs.LGstat.MLarXiv:1905.05393v12019
  44. Neighbourhood Consensus Networks

    Ignacio Rocco, Mircea Cimpoi, Relja Arandjelović +3

    cs.CVcs.LGarXiv:1810.10510v22018
  45. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

    Xiwei Hu, Rui Wang, Yixiao Fang +3

    cs.CVarXiv:2403.05135v12024
  46. Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations

    Leon Ranke, Wolfgang Hübner, Ronny Hug +2

    cs.CVcs.AIcs.LGarXiv:2608.27066v12026
  47. JCS: An Explainable COVID-19 Diagnosis System by Joint Classification and Segmentation

    Yu-Huan Wu, Shang-Hua Gao, Jie Mei +4

    eess.IVcs.CVcs.LGarXiv:2004.07054v32020
  48. Fast Point R-CNN

    Yilun Chen, Shu Liu, Xiaoyong Shen +1

    cs.CVarXiv:1908.02990v22019
  49. Scaling Open-Vocabulary Object Detection

    Matthias Minderer, Alexey Gritsenko, Neil Houlsby

    cs.CVarXiv:2306.09683v32023
  50. Hybrid Retrieval-Generation Reinforced Agent for Medical Image Report Generation

    Christy Y. Li, Xiaodan Liang, Zhiting Hu +1

    cs.CVarXiv:1805.08298v22018
  51. Data-Free Learning of Student Networks

    Hanting Chen, Yunhe Wang, Chang Xu +6

    cs.LGcs.CVstat.MLarXiv:1904.01186v42019
  52. Reasoning with Language Model Prompting: A Survey

    Shuofei Qiao, Yixin Ou, Ningyu Zhang +6

    cs.CLcs.AIcs.CVarXiv:2212.09597v82022
  53. SegDiff: Image Segmentation with Diffusion Probabilistic Models

    Tomer Amit, Tal Shaharbany, Eliya Nachmani +1

    cs.CVcs.AIcs.LGarXiv:2112.00390v32021
  54. CameraCtrl: Enabling Camera Control for Text-to-Video Generation

    Hao He, Yinghao Xu, Yuwei Guo +4

    cs.CVarXiv:2404.02101v22024
  55. Occlusion-aware R-CNN: Detecting Pedestrians in a Crowd

    Shifeng Zhang, Longyin Wen, Xiao Bian +2

    cs.CVarXiv:1807.08407v12018
  56. A Closed-form Solution to Photorealistic Image Stylization

    Yijun Li, Ming-Yu Liu, Xueting Li +2

    cs.CVarXiv:1802.06474v52018
  57. Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

    Haoning Wu, Erli Zhang, Liang Liao +6

    cs.CVcs.LGcs.MMarXiv:2211.04894v32022
  58. On the detection of synthetic images generated by diffusion models

    Riccardo Corvi, Davide Cozzolino, Giada Zingarini +3

    cs.CVarXiv:2211.00680v12022
  59. GuessWhat?! Visual object discovery through multi-modal dialogue

    Harm de Vries, Florian Strub, Sarath Chandar +3

    cs.AIcs.CVarXiv:1611.08481v22016
  60. Learning Less is More - 6D Camera Localization via 3D Surface Regression

    Eric Brachmann, Carsten Rother

    cs.CVarXiv:1711.10228v22017