Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,641 to 8,700 of 18,781

  1. AEGNN: Asynchronous Event-based Graph Neural Networks

    Simon Schaefer, Daniel Gehrig, Davide Scaramuzza

    cs.CVarXiv:2203.17149v32022
  2. MeGA-CDA: Memory Guided Attention for Category-Aware Unsupervised Domain Adaptive Object Detection

    Vibashan VS, Vikram Gupta, Poojan Oza +2

    cs.CVarXiv:2103.04224v22021
  3. Non-local Meets Global: An Iterative Paradigm for Hyperspectral Image Restoration

    Wei He, Quanming Yao, Chao Li +4

    eess.IVcs.CVarXiv:2010.12921v12020
  4. MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos

    Zhengqi Li, Richard Tucker, Forrester Cole +6

    cs.CVarXiv:2412.04463v22024
  5. DPGN: Distribution Propagation Graph Network for Few-shot Learning

    Ling Yang, Liangliang Li, Zilun Zhang +3

    cs.CVarXiv:2003.14247v22020
  6. NeRSemble: Multi-view Radiance Field Reconstruction of Human Heads

    Tobias Kirschstein, Shenhan Qian, Simon Giebenhain +2

    cs.CVarXiv:2305.03027v12023
  7. Novel View Synthesis of Dynamic Scenes with Globally Coherent Depths from a Monocular Camera

    Jae Shin Yoon, Kihwan Kim, Orazio Gallo +2

    cs.CVarXiv:2004.01294v12020
  8. VideoLLM-online: Online Video Large Language Model for Streaming Video

    Joya Chen, Zhaoyang Lv, Shiwei Wu +7

    cs.CVarXiv:2406.11816v12024
  9. DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models

    Chong Mou, Xintao Wang, Jiechong Song +2

    cs.CVarXiv:2307.02421v22023
  10. HPLFlowNet: Hierarchical Permutohedral Lattice FlowNet for Scene Flow Estimation on Large-scale Point Clouds

    Xiuye Gu, Yijie Wang, Chongruo wu +2

    cs.CVcs.LGeess.IVarXiv:1906.05332v12019
  11. Mask2Former for Video Instance Segmentation

    Bowen Cheng, Anwesa Choudhuri, Ishan Misra +3

    cs.CVcs.AIcs.LGarXiv:2112.10764v12021
  12. Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA

    Ronghang Hu, Amanpreet Singh, Trevor Darrell +1

    cs.CVcs.CLarXiv:1911.06258v32019
  13. 3D Point Cloud Processing and Learning for Autonomous Driving

    Siheng Chen, Baoan Liu, Chen Feng +2

    cs.CVeess.SParXiv:2003.00601v12020
  14. Decoders Matter for Semantic Segmentation: Data-Dependent Decoding Enables Flexible Feature Aggregation

    Zhi Tian, Tong He, Chunhua Shen +1

    cs.CVarXiv:1903.02120v32019
  15. Beyond Pixels: A Comprehensive Survey from Bottom-up to Semantic Image Segmentation and Cosegmentation

    Hongyuan Zhu, Fanman Meng, Jianfei Cai +1

    cs.CVarXiv:1502.00717v12015
  16. Real-time object detection method based on improved YOLOv4-tiny

    Zicong Jiang, Liquan Zhao, Shuaiyang Li +1

    cs.CVcs.AIarXiv:2011.04244v22020
  17. A Mutual Learning Method for Salient Object Detection with intertwined Multi-Supervision--Revised

    Runmin Wu, Mengyang Feng, Wenlong Guan +3

    cs.CVcs.AIarXiv:2509.21363v12025
  18. D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement

    Yansong Peng, Hebei Li, Peixi Wu +3

    cs.CVarXiv:2410.13842v12024
  19. Global Guidance Network for Breast Lesion Segmentation in Ultrasound Images

    Cheng Xue, Lei Zhu, Huazhu Fu +4

    eess.IVcs.CVcs.LGarXiv:2104.01896v12021
  20. Spatially-Aware Graph Neural Networks for Relational Behavior Forecasting from Sensor Data

    Sergio Casas, Cole Gulino, Renjie Liao +1

    cs.CVcs.LGcs.ROarXiv:1910.08233v12019
  21. Residual Conv-Deconv Grid Network for Semantic Segmentation

    Damien Fourure, Rémi Emonet, Elisa Fromont +3

    cs.CVarXiv:1707.07958v22017
  22. Learning to Balance Specificity and Invariance for In and Out of Domain Generalization

    Prithvijit Chattopadhyay, Yogesh Balaji, Judy Hoffman

    cs.CVcs.LGarXiv:2008.12839v12020
  23. Implicit Neural Representations for Image Compression

    Yannick Strümpler, Janis Postels, Ren Yang +2

    eess.IVcs.CVcs.LGarXiv:2112.04267v22021
  24. Mahotas: Open source software for scriptable computer vision

    Luis Pedro Coelho

    cs.CVcs.SEarXiv:1211.4907v22012
  25. Overcoming Catastrophic Forgetting with Unlabeled Data in the Wild

    Kibok Lee, Kimin Lee, Jinwoo Shin +1

    cs.CVcs.LGstat.MLarXiv:1903.12648v32019
  26. Adversarial Continual Learning

    Sayna Ebrahimi, Franziska Meier, Roberto Calandra +2

    cs.LGcs.AIcs.CVarXiv:2003.09553v22020
  27. HOPE-Net: A Graph-based Model for Hand-Object Pose Estimation

    Bardia Doosti, Shujon Naha, Majid Mirbagheri +1

    cs.CVarXiv:2004.00060v12020
  28. MODNet: Real-Time Trimap-Free Portrait Matting via Objective Decomposition

    Zhanghan Ke, Jiayu Sun, Kaican Li +2

    cs.CVarXiv:2011.11961v42020
  29. NeRO: Neural Geometry and BRDF Reconstruction of Reflective Objects from Multiview Images

    Yuan Liu, Peng Wang, Cheng Lin +5

    cs.CVcs.GRarXiv:2305.17398v12023
  30. 'Squeeze & Excite' Guided Few-Shot Segmentation of Volumetric Images

    Abhijit Guha Roy, Shayan Siddiqui, Sebastian Pölsterl +2

    cs.CVarXiv:1902.01314v22019
  31. PaliGemma 2: A Family of Versatile VLMs for Transfer

    Andreas Steiner, André Susano Pinto, Michael Tschannen +15

    cs.CVarXiv:2412.03555v12024
  32. FedMix: Approximation of Mixup under Mean Augmented Federated Learning

    Tehrim Yoon, Sumin Shin, Sung Ju Hwang +1

    cs.LGcs.AIcs.CVarXiv:2107.00233v12021
  33. Constructing Self-motivated Pyramid Curriculums for Cross-Domain Semantic Segmentation: A Non-Adversarial Approach

    Qing Lian, Fengmao Lv, Lixin Duan +1

    cs.CVarXiv:1908.09547v12019
  34. SeqTR: A Simple yet Universal Network for Visual Grounding

    Chaoyang Zhu, Yiyi Zhou, Yunhang Shen +7

    cs.CVarXiv:2203.16265v22022
  35. MatchFormer: Interleaving Attention in Transformers for Feature Matching

    Qing Wang, Jiaming Zhang, Kailun Yang +2

    cs.CVcs.ROeess.IVarXiv:2203.09645v32022
  36. A Poisson-Gaussian Denoising Dataset with Real Fluorescence Microscopy Images

    Yide Zhang, Yinhao Zhu, Evan Nichols +4

    cs.CVcs.LGeess.IVarXiv:1812.10366v22018
  37. TAO: A Large-Scale Benchmark for Tracking Any Object

    Achal Dave, Tarasha Khurana, Pavel Tokmakov +2

    cs.CVarXiv:2005.10356v12020
  38. Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational Learning

    Wentao Bao, Qi Yu, Yu Kong

    cs.CVarXiv:2008.00334v12020
  39. PHOCNet: A Deep Convolutional Neural Network for Word Spotting in Handwritten Documents

    Sebastian Sudholt, Gernot A. Fink

    cs.CVarXiv:1604.00187v32016
  40. StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models

    Zhizhong Wang, Lei Zhao, Wei Xing

    cs.CVarXiv:2308.07863v12023
  41. Deconfounded Video Moment Retrieval with Causal Intervention

    Xun Yang, Fuli Feng, Wei Ji +2

    cs.CVarXiv:2106.01534v12021
  42. Synthesis of Compositional Animations from Textual Descriptions

    Anindita Ghosh, Noshaba Cheema, Cennet Oguz +2

    cs.CVcs.LGarXiv:2103.14675v62021
  43. Fully Convolutional Architectures for Multi-Class Segmentation in Chest Radiographs

    Alexey A. Novikov, Dimitrios Lenis, David Major +3

    cs.CVcs.LGarXiv:1701.08816v42017
  44. Human Motion Prediction via Spatio-Temporal Inpainting

    Alejandro Hernandez Ruiz, Juergen Gall, Francesc Moreno-Noguer

    cs.CVarXiv:1812.05478v22018
  45. Towards a Joint Khmer Text Recognition and Word Segmentation

    Marry Kong, Rina Buoy, Sovisal Chenda +3

    cs.CVcs.CLarXiv:2608.30213v12026
  46. VIBE: Video Instruction-aligned Background music gEneration

    Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5

    cs.SDcs.AIcs.CLarXiv:2608.30125v12026
  47. Lot Machine: Multimodal Lot Extraction from Auction Catalogs

    Mathias Zinnen, Alisha Mund, Sabine Lang +3

    cs.CVcs.AIcs.CLarXiv:2608.30510v12026
  48. Activated Gradients for Deep Neural Networks

    Mei Liu, Liangming Chen, Xiaohao Du +2

    cs.CVcs.AIarXiv:2107.04228v12021
  49. Visual Representations for Semantic Target Driven Navigation

    Arsalan Mousavian, Alexander Toshev, Marek Fiser +3

    cs.CVarXiv:1805.06066v32018
  50. An Empirical Evaluation of Similarity Measures for Time Series Classification

    Joan Serrà, Josep Lluis Arcos

    cs.LGcs.CVstat.MLarXiv:1401.3973v12014
  51. Oriented Objects as pairs of Middle Lines

    Haoran Wei, Yue Zhang, Zhonghan Chang +3

    cs.CVarXiv:1912.10694v32019
  52. CRNet: Cross-Reference Networks for Few-Shot Segmentation

    Weide Liu, Chi Zhang, Guosheng Lin +1

    cs.CVarXiv:2003.10658v12020
  53. Snapshot Distillation: Teacher-Student Optimization in One Generation

    Chenglin Yang, Lingxi Xie, Chi Su +1

    cs.CVarXiv:1812.00123v12018
  54. Deep Roto-Translation Scattering for Object Classification

    Edouard Oyallon, Stéphane Mallat

    cs.CVarXiv:1412.8659v22014
  55. Deep Cascaded Bi-Network for Face Hallucination

    Shizhan Zhu, Sifei Liu, Chen Change Loy +1

    cs.CVarXiv:1607.05046v12016
  56. Deep Supervised Hashing with Triplet Labels

    Xiaofang Wang, Yi Shi, Kris M. Kitani

    cs.CVarXiv:1612.03900v12016
  57. VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

    Sicheng Xu, Guojun Chen, Yu-Xiao Guo +6

    cs.CVarXiv:2404.10667v22024
  58. Deep Imbalanced Attribute Classification using Visual Attention Aggregation

    Nikolaos Sarafianos, Xiang Xu, Ioannis A. Kakadiaris

    cs.CVarXiv:1807.03903v22018
  59. Invisible Steganography via Generative Adversarial Networks

    Ru Zhang, Shiqi Dong, Jianyi Liu

    cs.MMcs.CVarXiv:1807.08571v32018
  60. Deep Temporal Linear Encoding Networks

    Ali Diba, Vivek Sharma, Luc Van Gool

    cs.CVarXiv:1611.06678v12016