Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,841 to 12,900 of 18,822
Describing Multimedia Content using Attention-based Encoder--Decoder Networks
Kyunghyun Cho, Aaron Courville, Yoshua Bengio
cs.NEcs.CLcs.CVarXiv:1507.01053v12015KISS-GS: 3D Gaussian Splatting Compression Kept Simple
Wieland Morgenstern, Friedrich Elias Branschke, Florian Fleischmann +5
cs.CVarXiv:2608.26948v12026Behavior Transformers: Cloning $k$ modes with one stone
Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya +1
cs.LGcs.AIcs.CVarXiv:2206.11251v22022Understanding Deep Learning Techniques for Image Segmentation
Swarnendu Ghosh, Nibaran Das, Ishita Das +1
cs.CVcs.LGcs.NEarXiv:1907.06119v12019xView: Objects in Context in Overhead Imagery
Darius Lam, Richard Kuzma, Kevin McGee +5
cs.CVarXiv:1802.07856v12018Differentiable Jitter Correction using Deep Learning-based Image Quality Metric for Phase-Contrast Micro-CT
Junan Chen, Yiting Jia, Joscha Maier +6
cs.CVarXiv:2608.27034v12026Extended Object Tracking: Introduction, Overview and Applications
Karl Granstrom, Marcus Baum, Stephan Reuter
cs.CVeess.SPeess.SYarXiv:1604.00970v32016Recognize Anything: A Strong Image Tagging Model
Youcai Zhang, Xinyu Huang, Jinyu Ma +9
cs.CVarXiv:2306.03514v32023Real-time Action Recognition with Enhanced Motion Vector CNNs
Bowen Zhang, Limin Wang, Zhe Wang +2
cs.CVarXiv:1604.07669v12016Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
Xiaowei Zhou, Menglong Zhu, Spyridon Leonardos +2
cs.CVarXiv:1511.09439v22015Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
Xiu-Shen Wei, Jian-Hao Luo, Jianxin Wu +1
cs.CVarXiv:1604.04994v22016Advances in Medical Image Analysis with Vision Transformers: A Comprehensive Review
Reza Azad, Amirhossein Kazerouni, Moein Heidari +6
cs.CVarXiv:2301.03505v32023Combo Loss: Handling Input and Output Imbalance in Multi-Organ Segmentation
Saeid Asgari Taghanaki, Yefeng Zheng, S. Kevin Zhou +5
cs.CVarXiv:1805.02798v62018Virtual iEEG from Scalp EEG: Charting the Landscape of Source Imaging, Intracranial Inference and Reconstruction
Dongyi He, Xiangkai Wang, Hongjie Yan +3
cs.CVarXiv:2608.26998v12026UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed +3
cs.CVarXiv:2212.04497v32022Some Improvements on Deep Convolutional Neural Network Based Image Classification
Andrew G. Howard
cs.CVarXiv:1312.5402v12013Context-aware Synthesis for Video Frame Interpolation
Simon Niklaus, Feng Liu
cs.CVarXiv:1803.10967v12018Unified Concept Editing in Diffusion Models
Rohit Gandikota, Hadas Orgad, Yonatan Belinkov +2
cs.CVcs.LGarXiv:2308.14761v22023Vector Neurons: A General Framework for SO(3)-Equivariant Networks
Congyue Deng, Or Litany, Yueqi Duan +3
cs.CVarXiv:2104.12229v12021MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
Junyi Zhang, Charles Herrmann, Junhwa Hur +5
cs.CVarXiv:2410.03825v22024End-to-End Variational Networks for Accelerated MRI Reconstruction
Anuroop Sriram, Jure Zbontar, Tullie Murrell +5
eess.IVcs.CVarXiv:2004.06688v22020Multi-Scale Continuous CRFs as Sequential Deep Networks for Monocular Depth Estimation
Dan Xu, Elisa Ricci, Wanli Ouyang +2
cs.CVarXiv:1704.02157v12017FOSTER: Feature Boosting and Compression for Class-Incremental Learning
Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye +1
cs.CVcs.LGarXiv:2204.04662v22022DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion
Zixiang Zhao, Haowen Bai, Yuanzhi Zhu +7
cs.CVarXiv:2303.06840v22023Fruit recognition from images using deep learning
Horea Mureşan, Mihai Oltean
cs.CVarXiv:1712.00580v102017Learning Distilled Collaboration Graph for Multi-Agent Perception
Yiming Li, Shunli Ren, Pengxiang Wu +3
cs.CVcs.ROarXiv:2111.00643v22021SSH: Single Stage Headless Face Detector
Mahyar Najibi, Pouya Samangouei, Rama Chellappa +1
cs.CVarXiv:1708.03979v32017Weakly-Supervised Convolutional Neural Networks for Multimodal Image Registration
Yipeng Hu, Marc Modat, Eli Gibson +11
cs.CVcs.AIcs.LGarXiv:1807.03361v12018Deep-Plant: Plant Identification with convolutional neural networks
Sue Han Lee, Chee Seng Chan, Paul Wilkin +1
cs.CVcs.AIcs.NEarXiv:1506.08425v12015Deep representation learning for human motion prediction and classification
Judith Bütepage, Michael Black, Danica Kragic +1
cs.CVarXiv:1702.07486v22017Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model
Jiahao Li, Hao Tan, Kai Zhang +7
cs.CVarXiv:2311.06214v22023Understanding the Limitations of CNN-based Absolute Camera Pose Regression
Torsten Sattler, Qunjie Zhou, Marc Pollefeys +1
cs.CVarXiv:1903.07504v12019LRRNet: A Novel Representation Learning Guided Fusion Network for Infrared and Visible Images
Hui Li, Tianyang Xu, Xiao-Jun Wu +2
cs.CVarXiv:2304.05172v22023Semi-parametric Topological Memory for Navigation
Nikolay Savinov, Alexey Dosovitskiy, Vladlen Koltun
cs.LGcs.AIcs.CVarXiv:1803.00653v12018WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen +2
cs.CVcs.CLcs.IRarXiv:2103.01913v22021Beyond One-hot Encoding: lower dimensional target embedding
Pau Rodríguez, Miguel A. Bautista, Jordi Gonzàlez +1
cs.CVcs.AIarXiv:1806.10805v12018No New-Net
Fabian Isensee, Philipp Kickingereder, Wolfgang Wick +2
cs.CVarXiv:1809.10483v22018NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
Gengze Zhou, Yicong Hong, Qi Wu
cs.CVcs.AIcs.CLarXiv:2305.16986v32023Video Pixel Networks
Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan +4
cs.CVcs.LGarXiv:1610.00527v12016Video Processing from Electro-optical Sensors for Object Detection and Tracking in Maritime Environment: A Survey
D. K. Prasad, D. Rajan, L. Rachmawati +2
cs.CVarXiv:1611.05842v12016Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
Agniv Chatterjee, Georgios Pavlakos
cs.CVarXiv:2608.27407v12026BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Zhiliang Peng, Li Dong, Hangbo Bao +2
cs.CVarXiv:2208.06366v22022Population Based Augmentation: Efficient Learning of Augmentation Policy Schedules
Daniel Ho, Eric Liang, Ion Stoica +2
cs.CVcs.LGstat.MLarXiv:1905.05393v12019Neighbourhood Consensus Networks
Ignacio Rocco, Mircea Cimpoi, Relja Arandjelović +3
cs.CVcs.LGarXiv:1810.10510v22018ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Xiwei Hu, Rui Wang, Yixiao Fang +3
cs.CVarXiv:2403.05135v12024Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations
Leon Ranke, Wolfgang Hübner, Ronny Hug +2
cs.CVcs.AIcs.LGarXiv:2608.27066v12026JCS: An Explainable COVID-19 Diagnosis System by Joint Classification and Segmentation
Yu-Huan Wu, Shang-Hua Gao, Jie Mei +4
eess.IVcs.CVcs.LGarXiv:2004.07054v32020Fast Point R-CNN
Yilun Chen, Shu Liu, Xiaoyong Shen +1
cs.CVarXiv:1908.02990v22019Scaling Open-Vocabulary Object Detection
Matthias Minderer, Alexey Gritsenko, Neil Houlsby
cs.CVarXiv:2306.09683v32023Hybrid Retrieval-Generation Reinforced Agent for Medical Image Report Generation
Christy Y. Li, Xiaodan Liang, Zhiting Hu +1
cs.CVarXiv:1805.08298v22018Data-Free Learning of Student Networks
Hanting Chen, Yunhe Wang, Chang Xu +6
cs.LGcs.CVstat.MLarXiv:1904.01186v42019Reasoning with Language Model Prompting: A Survey
Shuofei Qiao, Yixin Ou, Ningyu Zhang +6
cs.CLcs.AIcs.CVarXiv:2212.09597v82022SegDiff: Image Segmentation with Diffusion Probabilistic Models
Tomer Amit, Tal Shaharbany, Eliya Nachmani +1
cs.CVcs.AIcs.LGarXiv:2112.00390v32021CameraCtrl: Enabling Camera Control for Text-to-Video Generation
Hao He, Yinghao Xu, Yuwei Guo +4
cs.CVarXiv:2404.02101v22024Occlusion-aware R-CNN: Detecting Pedestrians in a Crowd
Shifeng Zhang, Longyin Wen, Xiao Bian +2
cs.CVarXiv:1807.08407v12018A Closed-form Solution to Photorealistic Image Stylization
Yijun Li, Ming-Yu Liu, Xueting Li +2
cs.CVarXiv:1802.06474v52018Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives
Haoning Wu, Erli Zhang, Liang Liao +6
cs.CVcs.LGcs.MMarXiv:2211.04894v32022On the detection of synthetic images generated by diffusion models
Riccardo Corvi, Davide Cozzolino, Giada Zingarini +3
cs.CVarXiv:2211.00680v12022GuessWhat?! Visual object discovery through multi-modal dialogue
Harm de Vries, Florian Strub, Sarath Chandar +3
cs.AIcs.CVarXiv:1611.08481v22016Learning Less is More - 6D Camera Localization via 3D Surface Regression
Eric Brachmann, Carsten Rother
cs.CVarXiv:1711.10228v22017