Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,321 to 1,380 of 18,866

  1. Unsupervised Domain-Specific Deblurring via Disentangled Representations

    Boyu Lu, Jun-Cheng Chen, Rama Chellappa

    cs.CVarXiv:1903.01594v22019
  2. HiFaceGAN: Face Renovation via Collaborative Suppression and Replenishment

    Lingbo Yang, Chang Liu, Pan Wang +4

    cs.CVarXiv:2005.05005v22020
  3. MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

    Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat

    cs.CLcs.CVcs.MMarXiv:2609.11322v12026
  4. FilterReg: Robust and Efficient Probabilistic Point-Set Registration using Gaussian Filter and Twist Parameterization

    Wei Gao, Russ Tedrake

    cs.CVarXiv:1811.10136v32018
  5. RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free

    Cheng-Yang Fu, Mykhailo Shvets, Alexander C. Berg

    cs.CVarXiv:1901.03353v12019
  6. On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention

    Junyeop Lee, Sungrae Park, Jeonghun Baek +3

    cs.CVarXiv:1910.04396v12019
  7. Robust Mean Teacher for Continual and Gradual Test-Time Adaptation

    Mario Döbler, Robert A. Marsden, Bin Yang

    cs.CVarXiv:2211.13081v22022
  8. DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation

    Shentong Mo, Enze Xie, Ruihang Chu +4

    cs.CVcs.AIcs.LGarXiv:2307.01831v12023
  9. Wavelet-Based Dual-Branch Network for Image Demoireing

    Lin Liu, Jianzhuang Liu, Shanxin Yuan +4

    cs.CVarXiv:2007.07173v22020
  10. Structured Bird's-Eye-View Traffic Scene Understanding from Onboard Images

    Yigit Baran Can, Alexander Liniger, Danda Pani Paudel +1

    cs.CVarXiv:2110.01997v12021
  11. 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection

    He Wang, Yezhen Cong, Or Litany +2

    cs.CVarXiv:2012.04355v32020
  12. Deep Generative Modeling for Scene Synthesis via Hybrid Representations

    Zaiwei Zhang, Zhenpei Yang, Chongyang Ma +4

    cs.CVarXiv:1808.02084v12018
  13. Wireless End-to-End Image Transmission System using Semantic Communications

    Maheshi Lokumarambage, Vishnu Gowrisetty, Hossein Rezaei +3

    cs.CVarXiv:2302.13721v22023
  14. Motion-aware 3D Gaussian Splatting for Efficient Dynamic Scene Reconstruction

    Zhiyang Guo, Wengang Zhou, Li Li +2

    cs.CVarXiv:2403.11447v12024
  15. VideoBooth: Diffusion-based Video Generation with Image Prompts

    Yuming Jiang, Tianxing Wu, Shuai Yang +5

    cs.CVarXiv:2312.00777v12023
  16. OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

    Cong Wei, Zheyang Xiong, Weiming Ren +3

    cs.CVcs.AIarXiv:2411.07199v22024
  17. LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models

    Mingyang Xie, Numair Khan, Tianfu Wang +8

    cs.CVcs.LGarXiv:2601.14674v22026
  18. OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

    Jianjiang Yang, Peihang Li, Shanqing Xu +3

    cs.CLcs.CVarXiv:2609.11244v12026
  19. TableFormer: Table Structure Understanding with Transformers

    Ahmed Nassar, Nikolaos Livathinos, Maksym Lysak +1

    cs.CVcs.LGarXiv:2203.01017v22022
  20. Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting

    Syed Talal Wasim, Muzammal Naseer, Salman Khan +2

    cs.CVeess.IVarXiv:2304.03307v12023
  21. Efficient Semantic Video Segmentation with Per-frame Inference

    Yifan Liu, Chunhua Shen, Changqian Yu +1

    cs.CVarXiv:2002.11433v22020
  22. Detecting and Grounding Multi-Modal Media Manipulation

    Rui Shao, Tianxing Wu, Ziwei Liu

    cs.CVarXiv:2304.02556v12023
  23. HNeRV: A Hybrid Neural Representation for Videos

    Hao Chen, Matt Gwilliam, Ser-Nam Lim +1

    cs.CVarXiv:2304.02633v12023
  24. How Useful is Self-Supervised Pretraining for Visual Tasks?

    Alejandro Newell, Jia Deng

    cs.CVcs.LGarXiv:2003.14323v12020
  25. Structured Graph Learning for Clustering and Semi-supervised Classification

    Zhao Kang, Chong Peng, Qiang Cheng +4

    cs.LGcs.AIcs.CVarXiv:2008.13429v12020
  26. Self-supervised Knowledge Distillation Using Singular Value Decomposition

    Seung Hyun Lee, Dae Ha Kim, Byung Cheol Song

    cs.LGcs.CVstat.MLarXiv:1807.06819v12018
  27. MeshLRM: Large Reconstruction Model for High-Quality Meshes

    Xinyue Wei, Kai Zhang, Sai Bi +6

    cs.CVcs.GRarXiv:2404.12385v22024
  28. Time Series Change Point Detection with Self-Supervised Contrastive Predictive Coding

    Shohreh Deldari, Daniel V. Smith, Hao Xue +1

    cs.LGcs.AIcs.CVarXiv:2011.14097v52020
  29. GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Video

    Bruce X. B. Yu, Zhi Zhang, Yongxu Liu +3

    cs.CVarXiv:2307.05853v22023
  30. Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation

    Zifu Wan, Pingping Zhang, Yuhao Wang +4

    cs.CVarXiv:2404.04256v32024
  31. DMD: A Large-Scale Multi-Modal Driver Monitoring Dataset for Attention and Alertness Analysis

    Juan Diego Ortega, Neslihan Kose, Paola Cañas +5

    cs.CVcs.LGeess.IVarXiv:2008.12085v12020
  32. FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation

    Vladislav Bargatin, Alexander Yakovenko, Khaled Abud +1

    cs.CVarXiv:2609.11486v12026
  33. Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

    Zehan Wang, Haifeng Huang, Yang Zhao +2

    cs.CVarXiv:2308.08769v12023
  34. MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

    Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman +4

    cs.CVarXiv:2404.03413v12024
  35. Collaboration Helps Camera Overtake LiDAR in 3D Detection

    Yue Hu, Yifan Lu, Runsheng Xu +3

    cs.CVarXiv:2303.13560v12023
  36. SparseTrack: Multi-Object Tracking by Performing Scene Decomposition based on Pseudo-Depth

    Zelin Liu, Xinggang Wang, Cheng Wang +2

    cs.CVarXiv:2306.05238v22023
  37. MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation

    Jiaqi Chen, Bingqian Lin, Ran Xu +3

    cs.AIcs.CVcs.ROarXiv:2401.07314v32024
  38. Human Trajectory Prediction via Neural Social Physics

    Jiangbei Yue, Dinesh Manocha, He Wang

    cs.CVarXiv:2207.10435v22022
  39. Code-Aligned Autoencoders for Unsupervised Change Detection in Multimodal Remote Sensing Images

    Luigi T. Luppino, Mads A. Hansen, Michael Kampffmeyer +4

    cs.CVarXiv:2004.07011v12020
  40. Virchow: A Million-Slide Digital Pathology Foundation Model

    Eugene Vorontsov, Alican Bozkurt, Adam Casson +28

    eess.IVcs.CVcs.LGarXiv:2309.07778v62023
  41. Delving Deeper into Anti-aliasing in ConvNets

    Xueyan Zou, Fanyi Xiao, Zhiding Yu +1

    cs.CVarXiv:2008.09604v12020
  42. Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation

    Yunhao Gou, Kai Chen, Zhili Liu +6

    cs.CVarXiv:2403.09572v42024
  43. Human Motion Generation: A Survey

    Wentao Zhu, Xiaoxuan Ma, Dongwoo Ro +6

    cs.CVarXiv:2307.10894v32023
  44. The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth Estimation

    Saurabh Saxena, Charles Herrmann, Junhwa Hur +4

    cs.CVarXiv:2306.01923v22023
  45. Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation Vectors

    Ruihan Zhang, Prashan Madumal, Tim Miller +2

    cs.CVcs.AIcs.LGarXiv:2006.15417v42020
  46. Mi-Ripple: Restoring Images Degraded by Iterative AI Editing

    Jiayin Chen, Yicheng Xu, Muting Wang

    cs.CVarXiv:2609.11317v12026
  47. Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks

    Nan Wu, Stanisław Jastrzębski, Kyunghyun Cho +1

    cs.LGcs.CVarXiv:2202.05306v32022
  48. CORE: Consistent Representation Learning for Face Forgery Detection

    Yunsheng Ni, Depu Meng, Changqian Yu +3

    cs.CVcs.CRarXiv:2206.02749v12022
  49. Self-Supervised Video Forensics by Audio-Visual Anomaly Detection

    Chao Feng, Ziyang Chen, Andrew Owens

    cs.CVarXiv:2301.01767v22023
  50. Automatic Crack Detection on Road Pavements Using Encoder Decoder Architecture

    Zhun Fan, Chong Li, Ying Chen +4

    cs.CVcs.LGeess.IVarXiv:2007.00477v12020
  51. Deformable 3D Convolution for Video Super-Resolution

    Xinyi Ying, Longguang Wang, Yingqian Wang +3

    cs.CVarXiv:2004.02803v52020
  52. Efficient 3D Semantic Segmentation with Superpoint Transformer

    Damien Robert, Hugo Raguet, Loic Landrieu

    cs.CVarXiv:2306.08045v22023
  53. CMU-Net: A Strong ConvMixer-based Medical Ultrasound Image Segmentation Network

    Fenghe Tang, Lingtao Wang, Chunping Ning +2

    eess.IVcs.CVarXiv:2210.13012v42022
  54. Multivariate Confidence Calibration for Object Detection

    Fabian Küppers, Jan Kronenberger, Amirhossein Shantia +1

    cs.CVcs.LGstat.MLarXiv:2004.13546v12020
  55. Bilevel Layer-Positioning LoRA for Real Image Dehazing

    Yan Zhang, Long Ma, Yuxin Feng +3

    cs.CVarXiv:2603.10872v12026
  56. Using Text-to-Image Generation for Architectural Design Ideation

    Ville Paananen, Jonas Oppenlaender, Aku Visuri

    cs.HCcs.AIcs.CVarXiv:2304.10182v12023
  57. Pixel-GS: Density Control with Pixel-aware Gradient for 3D Gaussian Splatting

    Zheng Zhang, Wenbo Hu, Yixing Lao +2

    cs.CVarXiv:2403.15530v12024
  58. Surface Reconstruction from Point Clouds: A Survey and a Benchmark

    Zhangjin Huang, Yuxin Wen, Zihao Wang +2

    cs.CVarXiv:2205.02413v12022
  59. Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining

    Benjia Zhou, Zhigang Chen, Albert Clapés +5

    cs.CVarXiv:2307.14768v12023
  60. TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering

    Jingye Chen, Yupan Huang, Tengchao Lv +3

    cs.CVarXiv:2311.16465v12023