Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,561 to 7,620 of 18,841

  1. Plug-and-Play Unplugged: Optimization Free Reconstruction using Consensus Equilibrium

    Gregery T. Buzzard, Stanley H. Chan, Suhas Sreehari +1

    cs.CVmath.OCarXiv:1705.08983v32017
  2. Privacy-Preserving Human Activity Recognition from Extreme Low Resolution

    Michael S. Ryoo, Brandon Rothrock, Charles Fleming +1

    cs.CVarXiv:1604.03196v32016
  3. Talking Face Generation by Conditional Recurrent Adversarial Network

    Yang Song, Jingwen Zhu, Dawei Li +2

    cs.CVarXiv:1804.04786v32018
  4. Automatic segmenting teeth in X-ray images: Trends, a novel data set, benchmarking and future perspectives

    Gil Jader, Luciano Oliveira, Matheus Pithon

    cs.CVarXiv:1802.03086v12018
  5. Data-Efficient Networks for Multi-Contrast MRI Reconstruction based on a Generalized Content/Style Prior

    Chinmay Rao, Efe Ilıcak, Matthias J. P. van Osch +5

    eess.IVcs.CVarXiv:2609.01959v12026
  6. PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation

    Tianyuan Zhang, Hong-Xing Yu, Rundi Wu +5

    cs.CVcs.AIarXiv:2404.13026v22024
  7. Benchmarking Detection Transfer Learning with Vision Transformers

    Yanghao Li, Saining Xie, Xinlei Chen +3

    cs.CVarXiv:2111.11429v12021
  8. SAM3D: Segment Anything in 3D Scenes

    Yunhan Yang, Xiaoyang Wu, Tong He +2

    cs.CVarXiv:2306.03908v12023
  9. Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection

    Gongjie Zhang, Kaiwen Cui, Tzu-Yi Hung +1

    cs.CVarXiv:2103.15158v12021
  10. How Far is Video Generation from World Model: A Physical Law Perspective

    Bingyi Kang, Yang Yue, Rui Lu +5

    cs.CVcs.AIarXiv:2411.02385v22024
  11. How do neural networks see depth in single images?

    Tom van Dijk, Guido C. H. E. de Croon

    cs.CVcs.ROarXiv:1905.07005v12019
  12. Multi-Tool Image Editing Attribution in Facial Forgery

    Sheng Liu, Qiang Sheng, Danding Wang +3

    cs.CVcs.MMarXiv:2609.02751v12026
  13. Learning Predictive Representations for Deformable Objects Using Contrastive Estimation

    Wilson Yan, Ashwin Vangipuram, Pieter Abbeel +1

    cs.LGcs.CVcs.ROarXiv:2003.05436v12020
  14. The Devil is in Classification: A Simple Framework for Long-tail Object Detection and Instance Segmentation

    Tao Wang, Yu Li, Bingyi Kang +5

    cs.CVarXiv:2007.11978v52020
  15. Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer

    Zilong Huang, Youcheng Ben, Guozhong Luo +3

    cs.CVarXiv:2106.03650v12021
  16. HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model

    Di Wang, Meiqi Hu, Yao Jin +19

    cs.CVeess.IVarXiv:2406.11519v22024
  17. SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

    Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste +3

    cs.CVcs.LGarXiv:2206.07764v22022
  18. Rethinking Efficient Lane Detection via Curve Modeling

    Zhengyang Feng, Shaohua Guo, Xin Tan +3

    cs.CVcs.AIcs.LGarXiv:2203.02431v22022
  19. Score identity Distillation: Exponentially Fast Distillation of Pretrained Diffusion Models for One-Step Generation

    Mingyuan Zhou, Huangjie Zheng, Zhendong Wang +2

    cs.LGcs.AIcs.CVarXiv:2404.04057v32024
  20. Separating Style and Content for Generalized Style Transfer

    Yexun Zhang, Ya Zhang, Wenbin Cai +1

    cs.CVarXiv:1711.06454v62017
  21. Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation

    Yutong Liu, Nan Huang, Xu Cao +1

    cs.CVarXiv:2609.02864v12026
  22. Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges

    Qingyong Hu, Bo Yang, Sheikh Khalid +3

    cs.CVcs.AIcs.ROarXiv:2009.03137v32020
  23. Embedding Fourier for Ultra-High-Definition Low-Light Image Enhancement

    Chongyi Li, Chun-Le Guo, Man Zhou +4

    cs.CVarXiv:2302.11831v12023
  24. A Photometrically Calibrated Benchmark For Monocular Visual Odometry

    Jakob Engel, Vladyslav Usenko, Daniel Cremers

    cs.CVarXiv:1607.02555v22016
  25. MFAS: Multimodal Fusion Architecture Search

    Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux +2

    cs.LGcs.CVcs.NEarXiv:1903.06496v12019
  26. Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure

    Xiaohan Ding, Guiguang Ding, Yuchen Guo +1

    cs.LGcs.CVstat.MLarXiv:1904.03837v12019
  27. Morphology signal in whole slide image foundation models can automatically triage slides

    Ayushi Sinha, Shashank Yadav, Benjamin Holmes +9

    cs.CVcs.LGarXiv:2609.01987v12026
  28. A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting

    Junhao Zhuang, Yanhong Zeng, Wenran Liu +2

    cs.CVarXiv:2312.03594v42023
  29. Learning to Branch for Multi-Task Learning

    Pengsheng Guo, Chen-Yu Lee, Daniel Ulbricht

    cs.LGcs.CVstat.MLarXiv:2006.01895v22020
  30. Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input

    David Harwath, Adrià Recasens, Dídac Surís +3

    cs.CVcs.CLcs.SDarXiv:1804.01452v12018
  31. 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

    Sherwin Bahmani, Ivan Skorokhodov, Victor Rong +7

    cs.CVarXiv:2311.17984v22023
  32. A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

    Xianghong Fang, Wenlong Mou, Yuan Yuan +2

    cs.LGcs.CVarXiv:2609.02107v12026
  33. Plug-and-Play CNN for Crowd Motion Analysis: An Application in Abnormal Event Detection

    Mahdyar Ravanbakhsh, Moin Nabi, Hossein Mousavi +2

    cs.CVarXiv:1610.00307v32016
  34. Efficient Dense Modules of Asymmetric Convolution for Real-Time Semantic Segmentation

    Shao-Yuan Lo, Hsueh-Ming Hang, Sheng-Wei Chan +1

    cs.CVarXiv:1809.06323v32018
  35. A Closer Look at Local Aggregation Operators in Point Cloud Analysis

    Ze Liu, Han Hu, Yue Cao +2

    cs.CVcs.LGarXiv:2007.01294v12020
  36. Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection

    Chengjie Wang, Wenbing Zhu, Bin-Bin Gao +6

    cs.CVarXiv:2403.12580v12024
  37. Virtual Sparse Convolution for Multimodal 3D Object Detection

    Hai Wu, Chenglu Wen, Shaoshuai Shi +2

    cs.CVarXiv:2303.02314v12023
  38. Embedding Propagation: Smoother Manifold for Few-Shot Classification

    Pau Rodríguez, Issam Laradji, Alexandre Drouin +1

    cs.CVcs.LGarXiv:2003.04151v22020
  39. SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking

    Guangting Wang, Chong Luo, Zhiwei Xiong +1

    cs.CVarXiv:1904.04452v12019
  40. Foreground Segmentation Using a Triplet Convolutional Neural Network for Multiscale Feature Encoding

    Long Ang Lim, Hacer Yalim Keles

    cs.CVarXiv:1801.02225v12018
  41. SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

    Konstantin Klemmer, Esther Rolf, Caleb Robinson +2

    cs.CVcs.AIcs.CYarXiv:2311.17179v32023
  42. Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal Attention

    Byung-Hoon Kim, Jong Chul Ye, Jae-Jin Kim

    cs.CVcs.LGq-bio.NCarXiv:2105.13495v22021
  43. Taming Rectified Flow for Inversion and Editing

    Jiangshan Wang, Junfu Pu, Zhongang Qi +6

    cs.CVarXiv:2411.04746v32024
  44. Improved Bilinear Pooling with CNNs

    Tsung-Yu Lin, Subhransu Maji

    cs.CVarXiv:1707.06772v12017
  45. Actor and Action Video Segmentation from a Sentence

    Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li +1

    cs.CVarXiv:1803.07485v12018
  46. A Critic Evaluation of Methods for COVID-19 Automatic Detection from X-Ray Images

    Gianluca Maguolo, Loris Nanni

    eess.IVcs.CVcs.LGarXiv:2004.12823v42020
  47. FlowFusion: Dynamic Dense RGB-D SLAM Based on Optical Flow

    Tianwei Zhang, Huayan Zhang, Yang Li +2

    cs.ROcs.CVarXiv:2003.05102v12020
  48. Progressive Pseudo-Label Optimization for Point-Supervised Change Detection

    Hailong Ning, Hao Wang, Yimeng Wang +3

    cs.CVarXiv:2609.02171v12026
  49. Cycle Consistent Adversarial Denoising Network for Multiphase Coronary CT Angiography

    Eunhee Kang, Hyun Jung Koo, Dong Hyun Yang +2

    cs.CVcs.AIcs.LGarXiv:1806.09748v32018
  50. Diffusion-Encoding Gaussian Field for Joint k-q dMRI Reconstruction

    Zhibo Chen, Yajuan Huang, Yu Guan +3

    cs.CVarXiv:2609.02288v12026
  51. IM2CAD

    Hamid Izadinia, Qi Shan, Steven M. Seitz

    cs.CVarXiv:1608.05137v22016
  52. KiU-Net: Overcomplete Convolutional Architectures for Biomedical Image and Volumetric Segmentation

    Jeya Maria Jose Valanarasu, Vishwanath A. Sindagi, Ilker Hacihaliloglu +1

    eess.IVcs.CVarXiv:2010.01663v22020
  53. SINE: SINgle Image Editing with Text-to-Image Diffusion Models

    Zhixing Zhang, Ligong Han, Arnab Ghosh +2

    cs.CVcs.AIarXiv:2212.04489v22022
  54. Generative Adversarial Transformers

    Drew A. Hudson, C. Lawrence Zitnick

    cs.CVcs.AIcs.CLarXiv:2103.01209v42021
  55. Reducing the Memory Footprint of 3D Gaussian Splatting

    Panagiotis Papantonakis, Georgios Kopanas, Bernhard Kerbl +2

    cs.CVarXiv:2406.17074v12024
  56. Global Tracking Transformers

    Xingyi Zhou, Tianwei Yin, Vladlen Koltun +1

    cs.CVarXiv:2203.13250v22022
  57. Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition

    Liam Schoneveld, Alice Othmani, Hazem Abdelkawy

    cs.CVcs.LGcs.SDarXiv:2103.09154v22021
  58. Calibration and Comparative Analysis of Forward-Looking Sonar and 3D Sonar for Enhanced Underwater Object Recognition

    Aditya Penumarti, Khanh Dong, Zi-Hao Zhang +6

    cs.CVcs.ROarXiv:2608.29433v12026
  59. Fully Convolutional Network Ensembles for White Matter Hyperintensities Segmentation in MR Images

    Hongwei Li, Gongfa Jiang, Jianguo Zhang +4

    cs.CVarXiv:1802.05203v32018
  60. Uncertainty-Aware Multimodal Anti-UAV Detection via Evidential Fusion and Conflict-Discounted Belief Aggregation

    Sharanda Suttorp, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansour Alsahag

    cs.CVarXiv:2608.29235v12026