Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,321 to 7,380 of 18,927

  1. Cross-domain Face Presentation Attack Detection via Multi-domain Disentangled Representation Learning

    Guoqing Wang, Hu Han, Shiguang Shan +1

    cs.CVarXiv:2004.01959v12020
  2. Unsupervised Multi-Task Feature Learning on Point Clouds

    Kaveh Hassani, Mike Haley

    cs.CVcs.LGarXiv:1910.08207v12019
  3. PTQD: Accurate Post-Training Quantization for Diffusion Models

    Yefei He, Luping Liu, Jing Liu +3

    cs.CVarXiv:2305.10657v42023
  4. DRAMA: Joint Risk Localization and Captioning in Driving

    Srikanth Malla, Chiho Choi, Isht Dwivedi +2

    cs.CVcs.AIcs.LGarXiv:2209.10767v22022
  5. Conditional Deep Learning for Energy-Efficient and Enhanced Pattern Recognition

    Priyadarshini Panda, Abhronil Sengupta, Kaushik Roy

    cs.CVarXiv:1509.08971v62015
  6. Data-dependent Initializations of Convolutional Neural Networks

    Philipp Krähenbühl, Carl Doersch, Jeff Donahue +1

    cs.CVcs.LGarXiv:1511.06856v32015
  7. Learning a Rotation Invariant Detector with Rotatable Bounding Box

    Lei Liu, Zongxu Pan, Bin Lei

    cs.CVarXiv:1711.09405v12017
  8. Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolutions

    Ignacio Rocco, Relja Arandjelović, Josef Sivic

    cs.CVarXiv:2004.10566v12020
  9. Epipolar Transformers

    Yihui He, Rui Yan, Katerina Fragkiadaki +1

    cs.CVarXiv:2005.04551v12020
  10. FCN-Transformer Feature Fusion for Polyp Segmentation

    Edward Sanderson, Bogdan J. Matuszewski

    eess.IVcs.CVcs.LGarXiv:2208.08352v12022
  11. An Entropy-based Pruning Method for CNN Compression

    Jian-Hao Luo, Jianxin Wu

    cs.CVarXiv:1706.05791v12017
  12. Panoptic Lifting for 3D Scene Understanding with Neural Fields

    Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Buló +4

    cs.CVcs.LGarXiv:2212.09802v12022
  13. MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition

    Shuang Li, Kaixiong Gong, Chi Harold Liu +3

    cs.CVarXiv:2103.12579v32021
  14. Principia: Relational Physics Tests for Video Models

    Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan +1

    cs.CVarXiv:2609.04200v12026
  15. FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

    Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung

    cs.CVarXiv:2609.03563v12026
  16. Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors

    Xuanyi Dong, Shoou-I Yu, Xinshuo Weng +3

    cs.CVarXiv:1807.00966v22018
  17. What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention

    Antonino Furnari, Giovanni Maria Farinella

    cs.CVcs.AIarXiv:1905.09035v22019
  18. Hierarchical Long Short-Term Concurrent Memory for Human Interaction Recognition

    Xiangbo Shu, Jinhui Tang, Guo-Jun Qi +2

    cs.CVarXiv:1811.00270v12018
  19. Compressing AI Traffic: Standardized Neural Network Coding of Visual-Token Representations in Split Vision-Language Inference

    Reza Heidari, Hamed R. Tavakoli, Juho Kannala

    cs.CVeess.IVarXiv:2609.01200v12026
  20. Motion Guided Attention for Video Salient Object Detection

    Haofeng Li, Guanqi Chen, Guanbin Li +1

    cs.CVarXiv:1909.07061v22019
  21. BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net

    Marven Sherif, Amgad Elmasry, Youssef Ghazal +1

    cs.CVcs.AIarXiv:2609.01554v12026
  22. Grounded Video Description

    Luowei Zhou, Yannis Kalantidis, Xinlei Chen +2

    cs.CVarXiv:1812.06587v22018
  23. PointGPT: Auto-regressively Generative Pre-training from Point Clouds

    Guangyan Chen, Meiling Wang, Yi Yang +3

    cs.CVarXiv:2305.11487v22023
  24. WorldReward: Reward Modeling for Camera-Conditioned World Models

    Yibin Wang, Zehan Wang, Junshu Tang +13

    cs.CVarXiv:2609.03952v12026
  25. The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

    Yichen Liu, Quanwei Zhang, Haozhe Wang +7

    cs.MMcs.CVarXiv:2609.02367v12026
  26. CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

    Tingyu Song, Mingxin Li, Yanzhao Zhang +5

    cs.CVcs.AIcs.CLarXiv:2609.04083v12026
  27. GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions

    Junjie Wang, Jiemin Fang, Xiaopeng Zhang +2

    cs.CVcs.GRarXiv:2311.16037v22023
  28. Self-Supervised Learning for Cardiac MR Image Segmentation by Anatomical Position Prediction

    Wenjia Bai, Chen Chen, Giacomo Tarroni +6

    cs.CVarXiv:1907.02757v12019
  29. Sparse Representation-based Open Set Recognition

    He Zhang, Vishal M. Patel

    cs.CVarXiv:1705.02431v12017
  30. Fast Multi-class Dictionaries Learning with Geometrical Directions in MRI Reconstruction

    Zhifang Zhan, Jian-Feng Cai, Di Guo +3

    cs.CVmath.OCphysics.med-pharXiv:1503.02945v22015
  31. ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans

    Amirhosein Azarpour

    cs.CVcs.AIarXiv:2609.02333v12026
  32. LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

    Chuyan Chen, Haoxing Chen, Kun Chen +27

    cs.CVcs.AIarXiv:2609.03796v12026
  33. Improved Stereo Matching with Constant Highway Networks and Reflective Confidence Learning

    Amit Shaked, Lior Wolf

    cs.CVarXiv:1701.00165v12016
  34. Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks

    Avi Schwarzschild, Micah Goldblum, Arjun Gupta +2

    cs.LGcs.CRcs.CVarXiv:2006.12557v32020
  35. TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

    Liao Qu, Huichao Zhang, Yiheng Liu +7

    cs.CVcs.AIarXiv:2412.03069v22024
  36. World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models

    Chuhan Zhang, Seiji Ito, Kenta Hoshino +2

    cs.CVarXiv:2609.02159v12026
  37. SemanticAdv: Generating Adversarial Examples via Attribute-conditional Image Editing

    Haonan Qiu, Chaowei Xiao, Lei Yang +3

    cs.LGcs.CRcs.CVarXiv:1906.07927v42019
  38. HUGS: Human Gaussian Splats

    Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel +2

    cs.CVcs.GRarXiv:2311.17910v12023
  39. Editable Free-viewpoint Video Using a Layered Neural Representation

    Jiakai Zhang, Xinhang Liu, Xinyi Ye +6

    cs.CVcs.GRarXiv:2104.14786v12021
  40. Handwriting Trajectory Recovery via Autoregressive Ordered Stroke Instance Prediction

    En-Guang Wang, Yan-Ming Zhang, Fei Yin +1

    cs.CVarXiv:2609.02251v12026
  41. Debugging Tests for Model Explanations

    Julius Adebayo, Michael Muelly, Ilaria Liccardi +1

    cs.CVcs.LGarXiv:2011.05429v12020
  42. The Riemannian Geometry of Deep Generative Models

    Hang Shao, Abhishek Kumar, P. Thomas Fletcher

    cs.LGcs.CVstat.MLarXiv:1711.08014v12017
  43. Contrastive Learning for Image Captioning

    Bo Dai, Dahua Lin

    cs.CVarXiv:1710.02534v12017
  44. LoFi RADIO: A Distilled In-Domain Backbone Applied for Artifact-Severity Grading of Ultra-Low-Field Neonatal Brain MR

    Jonathan B. Martin, Yashwant Kurmi, Charlotte R. Sappo

    eess.IVcs.CVarXiv:2609.02676v12026
  45. Divide-and-Assemble: Learning Block-wise Memory for Unsupervised Anomaly Detection

    Jinlei Hou, Yingying Zhang, Qiaoyong Zhong +3

    cs.CVarXiv:2107.13118v12021
  46. Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts

    Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal +2

    cs.LGcs.AIcs.CVarXiv:2609.02101v12026
  47. Few-Shot Learning via Saliency-guided Hallucination of Samples

    Hongguang Zhang, Jing Zhang, Piotr Koniusz

    cs.CVarXiv:1904.03472v12019
  48. Distance Metric Learning using Graph Convolutional Networks: Application to Functional Brain Networks

    Sofia Ira Ktena, Sarah Parisot, Enzo Ferrante +4

    cs.CVcs.LGarXiv:1703.02161v22017
  49. Learning End-to-End Lossy Image Compression: A Benchmark

    Yueyu Hu, Wenhan Yang, Zhan Ma +1

    eess.IVcs.CVarXiv:2002.03711v42020
  50. Memory-Efficient Incremental Learning Through Feature Adaptation

    Ahmet Iscen, Jeffrey Zhang, Svetlana Lazebnik +1

    cs.CVarXiv:2004.00713v22020
  51. 3D Object Reconstruction from a Single Depth View with Adversarial Learning

    Bo Yang, Hongkai Wen, Sen Wang +3

    cs.CVcs.AIcs.LGarXiv:1708.07969v12017
  52. Conditional Channel Gated Networks for Task-Aware Continual Learning

    Davide Abati, Jakub Tomczak, Tijmen Blankevoort +3

    cs.CVcs.LGstat.MLarXiv:2004.00070v12020
  53. MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

    Ke Wang, Houxing Ren, Aojun Zhou +7

    cs.CLcs.AIcs.CVarXiv:2310.03731v12023
  54. MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

    Kaining Ying, Fanqing Meng, Jin Wang +19

    cs.CVarXiv:2404.16006v12024
  55. Exploiting Image-trained CNN Architectures for Unconstrained Video Classification

    Shengxin Zha, Florian Luisier, Walter Andrews +2

    cs.CVarXiv:1503.04144v32015
  56. Re-initialization Free Level Set Evolution via Reaction Diffusion

    Kaihua Zhang, Lei Zhang, Huihui Song +1

    cs.CVarXiv:1112.1496v32011
  57. Learning Progressive Modality-shared Transformers for Effective Visible-Infrared Person Re-identification

    Hu Lu, Xuezhang Zou, Pingping Zhang

    cs.CVcs.IRcs.MMarXiv:2212.00226v12022
  58. VanillaNet: the Power of Minimalism in Deep Learning

    Hanting Chen, Yunhe Wang, Jianyuan Guo +1

    cs.CVarXiv:2305.12972v22023
  59. RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation

    Titas Anciukevičius, Zexiang Xu, Matthew Fisher +4

    cs.CVcs.LGarXiv:2211.09869v42022
  60. Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics

    Yijie Lin, Ching-Chun Chang, Isao Echizen +2

    cs.CRcs.AIcs.CVarXiv:2609.02268v12026