Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,561 to 7,620 of 18,841
Plug-and-Play Unplugged: Optimization Free Reconstruction using Consensus Equilibrium
Gregery T. Buzzard, Stanley H. Chan, Suhas Sreehari +1
cs.CVmath.OCarXiv:1705.08983v32017Privacy-Preserving Human Activity Recognition from Extreme Low Resolution
Michael S. Ryoo, Brandon Rothrock, Charles Fleming +1
cs.CVarXiv:1604.03196v32016Talking Face Generation by Conditional Recurrent Adversarial Network
Yang Song, Jingwen Zhu, Dawei Li +2
cs.CVarXiv:1804.04786v32018Automatic segmenting teeth in X-ray images: Trends, a novel data set, benchmarking and future perspectives
Gil Jader, Luciano Oliveira, Matheus Pithon
cs.CVarXiv:1802.03086v12018Data-Efficient Networks for Multi-Contrast MRI Reconstruction based on a Generalized Content/Style Prior
Chinmay Rao, Efe Ilıcak, Matthias J. P. van Osch +5
eess.IVcs.CVarXiv:2609.01959v12026PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
Tianyuan Zhang, Hong-Xing Yu, Rundi Wu +5
cs.CVcs.AIarXiv:2404.13026v22024Benchmarking Detection Transfer Learning with Vision Transformers
Yanghao Li, Saining Xie, Xinlei Chen +3
cs.CVarXiv:2111.11429v12021SAM3D: Segment Anything in 3D Scenes
Yunhan Yang, Xiaoyang Wu, Tong He +2
cs.CVarXiv:2306.03908v12023Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection
Gongjie Zhang, Kaiwen Cui, Tzu-Yi Hung +1
cs.CVarXiv:2103.15158v12021How Far is Video Generation from World Model: A Physical Law Perspective
Bingyi Kang, Yang Yue, Rui Lu +5
cs.CVcs.AIarXiv:2411.02385v22024How do neural networks see depth in single images?
Tom van Dijk, Guido C. H. E. de Croon
cs.CVcs.ROarXiv:1905.07005v12019Multi-Tool Image Editing Attribution in Facial Forgery
Sheng Liu, Qiang Sheng, Danding Wang +3
cs.CVcs.MMarXiv:2609.02751v12026Learning Predictive Representations for Deformable Objects Using Contrastive Estimation
Wilson Yan, Ashwin Vangipuram, Pieter Abbeel +1
cs.LGcs.CVcs.ROarXiv:2003.05436v12020The Devil is in Classification: A Simple Framework for Long-tail Object Detection and Instance Segmentation
Tao Wang, Yu Li, Bingyi Kang +5
cs.CVarXiv:2007.11978v52020Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
Zilong Huang, Youcheng Ben, Guozhong Luo +3
cs.CVarXiv:2106.03650v12021HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model
Di Wang, Meiqi Hu, Yao Jin +19
cs.CVeess.IVarXiv:2406.11519v22024SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos
Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste +3
cs.CVcs.LGarXiv:2206.07764v22022Rethinking Efficient Lane Detection via Curve Modeling
Zhengyang Feng, Shaohua Guo, Xin Tan +3
cs.CVcs.AIcs.LGarXiv:2203.02431v22022Score identity Distillation: Exponentially Fast Distillation of Pretrained Diffusion Models for One-Step Generation
Mingyuan Zhou, Huangjie Zheng, Zhendong Wang +2
cs.LGcs.AIcs.CVarXiv:2404.04057v32024Separating Style and Content for Generalized Style Transfer
Yexun Zhang, Ya Zhang, Wenbin Cai +1
cs.CVarXiv:1711.06454v62017Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation
Yutong Liu, Nan Huang, Xu Cao +1
cs.CVarXiv:2609.02864v12026Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
Qingyong Hu, Bo Yang, Sheikh Khalid +3
cs.CVcs.AIcs.ROarXiv:2009.03137v32020Embedding Fourier for Ultra-High-Definition Low-Light Image Enhancement
Chongyi Li, Chun-Le Guo, Man Zhou +4
cs.CVarXiv:2302.11831v12023A Photometrically Calibrated Benchmark For Monocular Visual Odometry
Jakob Engel, Vladyslav Usenko, Daniel Cremers
cs.CVarXiv:1607.02555v22016MFAS: Multimodal Fusion Architecture Search
Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux +2
cs.LGcs.CVcs.NEarXiv:1903.06496v12019Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
Xiaohan Ding, Guiguang Ding, Yuchen Guo +1
cs.LGcs.CVstat.MLarXiv:1904.03837v12019Morphology signal in whole slide image foundation models can automatically triage slides
Ayushi Sinha, Shashank Yadav, Benjamin Holmes +9
cs.CVcs.LGarXiv:2609.01987v12026A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
Junhao Zhuang, Yanhong Zeng, Wenran Liu +2
cs.CVarXiv:2312.03594v42023Learning to Branch for Multi-Task Learning
Pengsheng Guo, Chen-Yu Lee, Daniel Ulbricht
cs.LGcs.CVstat.MLarXiv:2006.01895v22020Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input
David Harwath, Adrià Recasens, Dídac Surís +3
cs.CVcs.CLcs.SDarXiv:1804.01452v120184D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling
Sherwin Bahmani, Ivan Skorokhodov, Victor Rong +7
cs.CVarXiv:2311.17984v22023A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization
Xianghong Fang, Wenlong Mou, Yuan Yuan +2
cs.LGcs.CVarXiv:2609.02107v12026Plug-and-Play CNN for Crowd Motion Analysis: An Application in Abnormal Event Detection
Mahdyar Ravanbakhsh, Moin Nabi, Hossein Mousavi +2
cs.CVarXiv:1610.00307v32016Efficient Dense Modules of Asymmetric Convolution for Real-Time Semantic Segmentation
Shao-Yuan Lo, Hsueh-Ming Hang, Sheng-Wei Chan +1
cs.CVarXiv:1809.06323v32018A Closer Look at Local Aggregation Operators in Point Cloud Analysis
Ze Liu, Han Hu, Yue Cao +2
cs.CVcs.LGarXiv:2007.01294v12020Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection
Chengjie Wang, Wenbing Zhu, Bin-Bin Gao +6
cs.CVarXiv:2403.12580v12024Virtual Sparse Convolution for Multimodal 3D Object Detection
Hai Wu, Chenglu Wen, Shaoshuai Shi +2
cs.CVarXiv:2303.02314v12023Embedding Propagation: Smoother Manifold for Few-Shot Classification
Pau Rodríguez, Issam Laradji, Alexandre Drouin +1
cs.CVcs.LGarXiv:2003.04151v22020SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking
Guangting Wang, Chong Luo, Zhiwei Xiong +1
cs.CVarXiv:1904.04452v12019Foreground Segmentation Using a Triplet Convolutional Neural Network for Multiscale Feature Encoding
Long Ang Lim, Hacer Yalim Keles
cs.CVarXiv:1801.02225v12018SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery
Konstantin Klemmer, Esther Rolf, Caleb Robinson +2
cs.CVcs.AIcs.CYarXiv:2311.17179v32023Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal Attention
Byung-Hoon Kim, Jong Chul Ye, Jae-Jin Kim
cs.CVcs.LGq-bio.NCarXiv:2105.13495v22021Taming Rectified Flow for Inversion and Editing
Jiangshan Wang, Junfu Pu, Zhongang Qi +6
cs.CVarXiv:2411.04746v32024Improved Bilinear Pooling with CNNs
Tsung-Yu Lin, Subhransu Maji
cs.CVarXiv:1707.06772v12017Actor and Action Video Segmentation from a Sentence
Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li +1
cs.CVarXiv:1803.07485v12018A Critic Evaluation of Methods for COVID-19 Automatic Detection from X-Ray Images
Gianluca Maguolo, Loris Nanni
eess.IVcs.CVcs.LGarXiv:2004.12823v42020FlowFusion: Dynamic Dense RGB-D SLAM Based on Optical Flow
Tianwei Zhang, Huayan Zhang, Yang Li +2
cs.ROcs.CVarXiv:2003.05102v12020Progressive Pseudo-Label Optimization for Point-Supervised Change Detection
Hailong Ning, Hao Wang, Yimeng Wang +3
cs.CVarXiv:2609.02171v12026Cycle Consistent Adversarial Denoising Network for Multiphase Coronary CT Angiography
Eunhee Kang, Hyun Jung Koo, Dong Hyun Yang +2
cs.CVcs.AIcs.LGarXiv:1806.09748v32018Diffusion-Encoding Gaussian Field for Joint k-q dMRI Reconstruction
Zhibo Chen, Yajuan Huang, Yu Guan +3
cs.CVarXiv:2609.02288v12026IM2CAD
Hamid Izadinia, Qi Shan, Steven M. Seitz
cs.CVarXiv:1608.05137v22016KiU-Net: Overcomplete Convolutional Architectures for Biomedical Image and Volumetric Segmentation
Jeya Maria Jose Valanarasu, Vishwanath A. Sindagi, Ilker Hacihaliloglu +1
eess.IVcs.CVarXiv:2010.01663v22020SINE: SINgle Image Editing with Text-to-Image Diffusion Models
Zhixing Zhang, Ligong Han, Arnab Ghosh +2
cs.CVcs.AIarXiv:2212.04489v22022Generative Adversarial Transformers
Drew A. Hudson, C. Lawrence Zitnick
cs.CVcs.AIcs.CLarXiv:2103.01209v42021Reducing the Memory Footprint of 3D Gaussian Splatting
Panagiotis Papantonakis, Georgios Kopanas, Bernhard Kerbl +2
cs.CVarXiv:2406.17074v12024Global Tracking Transformers
Xingyi Zhou, Tianwei Yin, Vladlen Koltun +1
cs.CVarXiv:2203.13250v22022Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition
Liam Schoneveld, Alice Othmani, Hazem Abdelkawy
cs.CVcs.LGcs.SDarXiv:2103.09154v22021Calibration and Comparative Analysis of Forward-Looking Sonar and 3D Sonar for Enhanced Underwater Object Recognition
Aditya Penumarti, Khanh Dong, Zi-Hao Zhang +6
cs.CVcs.ROarXiv:2608.29433v12026Fully Convolutional Network Ensembles for White Matter Hyperintensities Segmentation in MR Images
Hongwei Li, Gongfa Jiang, Jianguo Zhang +4
cs.CVarXiv:1802.05203v32018Uncertainty-Aware Multimodal Anti-UAV Detection via Evidential Fusion and Conflict-Discounted Belief Aggregation
Sharanda Suttorp, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansour Alsahag
cs.CVarXiv:2608.29235v12026