Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,561 to 1,620 of 18,779
MoVQ: Modulating Quantized Vectors for High-Fidelity Image Generation
Chuanxia Zheng, Long Tung Vuong, Jianfei Cai +1
cs.CVarXiv:2209.09002v12022Active Transfer Learning Network: A Unified Deep Joint Spectral-Spatial Feature Learning Model For Hyperspectral Image Classification
Cheng Deng, Yumeng Xue, Xianglong Liu +2
cs.CVarXiv:1904.02454v12019Learning Blind Motion Deblurring
Patrick Wieschollek, Michael Hirsch, Bernhard Schölkopf +1
cs.CVarXiv:1708.04208v12017Pedestrian Path, Pose and Intention Prediction through Gaussian Process Dynamical Models and Pedestrian Activity Recognition
Raul Quintero, Ignacio Parra, David Fernandez Llorca +1
cs.CVarXiv:2004.14747v12020A deep learning approach to detecting volcano deformation from satellite imagery using synthetic datasets
Nantheera Anantrasirichai, Juliet Biggs, Fabien Albino +1
cs.CVeess.IVarXiv:1905.07286v12019COUCH: Towards Controllable Human-Chair Interactions
Xiaohan Zhang, Bharat Lal Bhatnagar, Vladimir Guzov +2
cs.CVarXiv:2205.00541v12022Visual Search at Pinterest
Yushi Jing, David Liu, Dmitry Kislyuk +4
cs.CVarXiv:1505.07647v32015Deep Learning in Diabetic Foot Ulcers Detection: A Comprehensive Evaluation
Moi Hoon Yap, Ryo Hachiuma, Azadeh Alavi +18
cs.CVarXiv:2010.03341v32020LightMedSeg-ISLES: Stroke Lesion Segmentation with 81x Fewer Parameters than nnU-Net
Giorgi Nikvashvili, Hanxue Gu, Jie Bao +2
cs.CVcs.LGarXiv:2609.09634v12026Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions
Yuxing Long, Xiaoqi Li, Wenzhe Cai +1
cs.ROcs.AIcs.CLarXiv:2309.11382v12023Probabilistic Monocular 3D Human Pose Estimation with Normalizing Flows
Tom Wehrbein, Marco Rudolph, Bodo Rosenhahn +1
cs.CVarXiv:2107.13788v22021ZigMa: A DiT-style Zigzag Mamba Diffusion Model
Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui +4
cs.CVcs.AIcs.CLarXiv:2403.13802v32024MethaneFuse: Learning from Multi-Sensor Satellite Observations for Methane Plume Detection
Yuyao Wang, Juliana Y. Leung, Di Niu
cs.CVcs.LGarXiv:2609.09762v12026LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis
Zejiang Shen, Ruochen Zhang, Melissa Dell +3
cs.CVcs.AIarXiv:2103.15348v22021Meta-Learning with Task-Adaptive Loss Function for Few-Shot Learning
Sungyong Baik, Janghoon Choi, Heewon Kim +3
cs.LGcs.CVarXiv:2110.03909v22021Domain Generalization with Domain-Specific Aggregation Modules
Antonio D'Innocente, Barbara Caputo
cs.CVarXiv:1809.10966v12018Deep Hyperspherical Learning
Weiyang Liu, Yan-Ming Zhang, Xingguo Li +4
cs.LGcs.CVstat.MLarXiv:1711.03189v52017Fine-Grained Car Detection for Visual Census Estimation
Timnit Gebru, Jonathan Krause, Yilun Wang +3
cs.CVarXiv:1709.02480v12017TAC-GAN - Text Conditioned Auxiliary Classifier Generative Adversarial Network
Ayushman Dash, John Cristian Borges Gamboa, Sheraz Ahmed +2
cs.CVarXiv:1703.06412v22017A General Pipeline for 3D Detection of Vehicles
Xinxin Du, Marcelo H. Ang, Sertac Karaman +1
cs.CVeess.IVstat.MLarXiv:1803.00387v120182D Car Detection in Radar Data with PointNets
Andreas Danzer, Thomas Griebel, Martin Bach +1
cs.CVcs.LGstat.MLarXiv:1904.08414v32019D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions
Sammy Christen, Muhammed Kocabas, Emre Aksan +3
cs.CVcs.LGcs.ROarXiv:2112.03028v22021Top-Down Feedback for Crowd Counting Convolutional Neural Network
Deepak Babu Sam, R. Venkatesh Babu
cs.CVarXiv:1807.08881v22018Learning Disentangled Semantic Representation for Domain Adaptation
Ruichu Cai, Zijian Li, Pengfei Wei +3
cs.CVcs.LGarXiv:2012.11807v12020Retrieval Augmented Classification for Long-Tail Visual Recognition
Alexander Long, Wei Yin, Thalaiyasingam Ajanthan +6
cs.CVarXiv:2202.11233v12022Face Anti-Spoofing with Human Material Perception
Zitong Yu, Xiaobai Li, Xuesong Niu +2
cs.CVarXiv:2007.02157v12020DeepTravel: a Neural Network Based Travel Time Estimation Model with Auxiliary Supervision
Hanyuan Zhang, Hao Wu, Weiwei Sun +1
cs.LGcs.CVstat.MLarXiv:1802.02147v12018End-to-end Trained CNN Encode-Decoder Networks for Image Steganography
Atique ur Rehman, Rafia Rahim, M Shahroz Nadeem +1
cs.MMcs.CVarXiv:1711.07201v12017ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
Guanxing Lu, Shiyi Zhang, Ziwei Wang +3
cs.ROcs.CVarXiv:2403.08321v22024From ImageNet to Image Classification: Contextualizing Progress on Benchmarks
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom +2
cs.CVcs.LGstat.MLarXiv:2005.11295v12020LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour Segmentation
Yi Luo, Yike Guo, Wenxuan Li +3
cs.CVcs.LGarXiv:2609.09477v12026All-In-One Underwater Image Enhancement using Domain-Adversarial Learning
Pritish Uplavikar, Zhenyu Wu, Zhangyang Wang
cs.CVarXiv:1905.13342v12019We are More than Our Joints: Predicting how 3D Bodies Move
Yan Zhang, Michael J. Black, Siyu Tang
cs.CVarXiv:2012.00619v22020Joint super-resolution and synthesis of 1 mm isotropic MP-RAGE volumes from clinical MRI exams with scans of different orientation, resolution and contrast
Juan Eugenio Iglesias, Benjamin Billot, Yael Balbastre +6
eess.IVcs.CVarXiv:2012.13340v12020Auxiliary Signal-Guided Knowledge Encoder-Decoder for Medical Report Generation
Mingjie Li, Fuyu Wang, Xiaojun Chang +1
cs.CVcs.CLeess.IVarXiv:2006.03744v120203D Human Pose Estimation via Intuitive Physics
Shashank Tripathi, Lea Müller, Chun-Hao P. Huang +3
cs.CVcs.AIcs.GRarXiv:2303.18246v32023KING: Generating Safety-Critical Driving Scenarios for Robust Imitation via Kinematics Gradients
Niklas Hanselmann, Katrin Renz, Kashyap Chitta +2
cs.ROcs.CVcs.LGarXiv:2204.13683v12022Space-time Mixing Attention for Video Transformer
Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2
cs.CVcs.AIcs.LGarXiv:2106.05968v22021TINYCD: A (Not So) Deep Learning Model For Change Detection
Andrea Codegoni, Gabriele Lombardi, Alessandro Ferrari
cs.CVcs.LGeess.IVarXiv:2207.13159v22022Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction
Bo Yang, Sen Wang, Andrew Markham +1
cs.CVcs.AIcs.LGarXiv:1808.00758v22018Robot Navigation in Crowds by Graph Convolutional Networks with Attention Learned from Human Gaze
Yuying Chen, Congcong Liu, Ming Liu +1
cs.ROcs.AIcs.CVarXiv:1909.10400v12019VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
Jiazheng Xu, Yu Huang, Jiale Cheng +19
cs.CVarXiv:2412.21059v42024DeepDRR -- A Catalyst for Machine Learning in Fluoroscopy-guided Procedures
Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee +4
physics.med-phcs.CVarXiv:1803.08606v12018StopThePop: Sorted Gaussian Splatting for View-Consistent Real-time Rendering
Lukas Radl, Michael Steiner, Mathias Parger +3
cs.GRcs.CVarXiv:2402.00525v32024Training CNNs with Low-Rank Filters for Efficient Image Classification
Yani Ioannou, Duncan Robertson, Jamie Shotton +2
cs.CVcs.LGcs.NEarXiv:1511.06744v32015Face Recognition Using Deep Multi-Pose Representations
Wael AbdAlmageed, Yue Wua, Stephen Rawlsa +9
cs.CVarXiv:1603.07388v12016PreDiff: Precipitation Nowcasting with Latent Diffusion Models
Zhihan Gao, Xingjian Shi, Boran Han +6
cs.LGcs.AIcs.CVarXiv:2307.10422v22023CubiCasa5K: A Dataset and an Improved Multi-Task Model for Floorplan Image Analysis
Ahti Kalervo, Juha Ylioinas, Markus Häikiö +2
cs.CVarXiv:1904.01920v12019TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4
cs.CVarXiv:2104.08271v22021A multilevel thresholding algorithm using Electromagnetism Optimization
Diego Oliva, Erik Cuevas, Gonzalo Pajares +2
cs.CVarXiv:1406.6336v12014Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation
Bingchen Zhao, Kai Han
cs.CVarXiv:2107.03358v22021MSeg3D: Multi-modal 3D Semantic Segmentation for Autonomous Driving
Jiale Li, Hang Dai, Hao Han +1
cs.CVarXiv:2303.08600v12023Beyond Physical Connections: Tree Models in Human Pose Estimation
Fang Wang, Yi Li
cs.CVarXiv:1305.2269v12013Multi-Angle Point Cloud-VAE: Unsupervised Feature Learning for 3D Point Clouds from Multiple Angles by Joint Self-Reconstruction and Half-to-Half Prediction
Zhizhong Han, Xiyang Wang, Yu-Shen Liu +1
cs.CVarXiv:1907.12704v12019Planar Prior Assisted PatchMatch Multi-View Stereo
Qingshan Xu, Wenbing Tao
cs.CVarXiv:1912.11744v12019Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta +1
cs.ROcs.CVarXiv:2405.01527v22024A Closer Look at the Explainability of Contrastive Language-Image Pre-training
Yi Li, Hualiang Wang, Yiqun Duan +2
cs.CVarXiv:2304.05653v22023Total Denoising: Unsupervised Learning of 3D Point Cloud Cleaning
Pedro Hermosilla, Tobias Ritschel, Timo Ropinski
cs.CVcs.GRarXiv:1904.07615v22019Analyzing and Mitigating the Impact of Permanent Faults on a Systolic Array Based Neural Network Accelerator
Jeff Zhang, Tianyu Gu, Kanad Basu +1
cs.LGcs.ARcs.CVarXiv:1802.04657v22018VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
Xiang Li, Jian Ding, Mohamed Elhoseiny
cs.CVarXiv:2406.12384v22024