Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,781 to 12,840 of 18,830

  1. Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS

    Kangmin Seo, Jae-Pil Heo

    cs.CVcs.AIarXiv:2608.26951v12026
  2. face anti-spoofing based on color texture analysis

    Zinelabidine Boulkenafet, Jukka Komulainen, Abdenour Hadid

    cs.CVarXiv:1511.06316v12015
  3. SoftTriple Loss: Deep Metric Learning Without Triplet Sampling

    Qi Qian, Lei Shang, Baigui Sun +3

    cs.CVarXiv:1909.05235v22019
  4. VRT: A Video Restoration Transformer

    Jingyun Liang, Jiezhang Cao, Yuchen Fan +5

    cs.CVeess.IVarXiv:2201.12288v22022
  5. Reinforced Continual Learning

    Ju Xu, Zhanxing Zhu

    cs.LGcs.CVstat.MLarXiv:1805.12369v12018
  6. DCSAU-Net: A Deeper and More Compact Split-Attention U-Net for Medical Image Segmentation

    Qing Xu, Zhicheng Ma, Na HE +1

    eess.IVcs.AIcs.CVarXiv:2202.00972v22022
  7. SSD: A Unified Framework for Self-Supervised Outlier Detection

    Vikash Sehwag, Mung Chiang, Prateek Mittal

    cs.CVcs.AIcs.LGarXiv:2103.12051v12021
  8. Learnable Triangulation of Human Pose

    Karim Iskakov, Egor Burkov, Victor Lempitsky +1

    cs.CVcs.AIarXiv:1905.05754v12019
    Summaries:한국어
  9. Masked Siamese Networks for Label-Efficient Learning

    Mahmoud Assran, Mathilde Caron, Ishan Misra +6

    cs.LGcs.AIcs.CVarXiv:2204.07141v12022
  10. ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction

    Bojun Zhang, Junhong Liang, Feifei Zhai +2

    cs.CVarXiv:2608.27154v12026
  11. PPF-FoldNet: Unsupervised Learning of Rotation Invariant 3D Local Descriptors

    Haowen Deng, Tolga Birdal, Slobodan Ilic

    cs.CVcs.CGcs.LGarXiv:1808.10322v12018
  12. Rank Minimization for Snapshot Compressive Imaging

    Yang Liu, Xin Yuan, Jinli Suo +2

    cs.CVarXiv:1807.07837v12018
  13. Scaling Language-Image Pre-training via Masking

    Yanghao Li, Haoqi Fan, Ronghang Hu +2

    cs.CVarXiv:2212.00794v22022
    Summaries:한국어
  14. VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation

    Zhengxiong Luo, Dayou Chen, Yingya Zhang +6

    cs.CVarXiv:2303.08320v42023
  15. MaskFusion: Real-Time Recognition, Tracking and Reconstruction of Multiple Moving Objects

    Martin Rünz, Maud Buffier, Lourdes Agapito

    cs.CVcs.ROarXiv:1804.09194v22018
  16. A Model-driven Deep Neural Network for Single Image Rain Removal

    Hong Wang, Qi Xie, Qian Zhao +1

    eess.IVcs.CVarXiv:2005.01333v12020
  17. The Segment Anything Model (SAM) for Remote Sensing Applications: From Zero to One Shot

    Lucas Prado Osco, Qiusheng Wu, Eduardo Lopes de Lemos +4

    cs.CVarXiv:2306.16623v22023
  18. Data-efficient crack quantification in lithium-ion cathodes using foundation model transfer

    Thorsten Tegetmeyer-Kleine, Thomas Schmitt, Phillip Aquino +3

    cond-mat.mtrl-scics.CVcs.LGarXiv:2608.27162v12026
  19. Burst Denoising with Kernel Prediction Networks

    Ben Mildenhall, Jonathan T. Barron, Jiawen Chen +3

    cs.CVarXiv:1712.02327v22017
  20. From Motion Blur to Motion Flow: a Deep Learning Solution for Removing Heterogeneous Motion Blur

    Dong Gong, Jie Yang, Lingqiao Liu +5

    cs.CVarXiv:1612.02583v12016
  21. Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning

    Lin Zhang, Li Shen, Liang Ding +2

    cs.LGcs.CVarXiv:2203.09249v22022
  22. Remote Photoplethysmograph Signal Measurement from Facial Videos Using Spatio-Temporal Networks

    Zitong Yu, Xiaobai Li, Guoying Zhao

    cs.CVarXiv:1905.02419v22019
  23. LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

    Lukas Kuhn, Lucas Maes, Giuseppe Serra +4

    cs.CVcs.AIarXiv:2608.27395v12026
  24. SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous Driving

    Yi Wei, Linqing Zhao, Wenzhao Zheng +3

    cs.CVarXiv:2303.09551v22023
  25. IQA: Visual Question Answering in Interactive Environments

    Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari +3

    cs.CVarXiv:1712.03316v32017
  26. Splatter Image: Ultra-Fast Single-View 3D Reconstruction

    Stanislaw Szymanowicz, Christian Rupprecht, Andrea Vedaldi

    cs.CVarXiv:2312.13150v22023
  27. Real-time Semantic Segmentation of Crop and Weed for Precision Agriculture Robots Leveraging Background Knowledge in CNNs

    Andres Milioto, Philipp Lottes, Cyrill Stachniss

    cs.CVcs.ROarXiv:1709.06764v22017
  28. Few-Shot Adversarial Domain Adaptation

    Saeid Motiian, Quinn Jones, Seyed Mehdi Iranmanesh +1

    cs.CVarXiv:1711.02536v12017
  29. Simple Open-Vocabulary Object Detection with Vision Transformers

    Matthias Minderer, Alexey Gritsenko, Austin Stone +11

    cs.CVarXiv:2205.06230v22022
  30. Decoupled I/O-Dominant Pipelines for Large-Scale Whole-Slide Image Embedding Extraction

    Mayanka Chandrashekar, Xi Zhang, Ethan Seefried +3

    cs.DCcs.CVarXiv:2608.27278v12026
  31. Vision-centric generative AI models: A software-hardware perspective

    Eleni Tselepi, Cristian Sestito, Shady Agwa +1

    cs.CVcs.ARarXiv:2608.27199v12026
  32. Adaptive Neural Networks for Efficient Inference

    Tolga Bolukbasi, Joseph Wang, Ofer Dekel +1

    cs.LGcs.CVcs.NEarXiv:1702.07811v22017
  33. ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification

    M. Middleton, H. Kayan, B. Sen Bhattacharya +7

    cs.NEcs.AIcs.CVarXiv:2608.27150v12026
  34. CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

    Kechen Liu, Ola Shorinwa

    cs.ROcs.AIcs.CVarXiv:2608.27406v12026
  35. Constructing Stronger and Faster Baselines for Skeleton-based Action Recognition

    Yi-Fan Song, Zhang Zhang, Caifeng Shan +1

    cs.CVarXiv:2106.15125v22021
  36. Disentangling Physical Dynamics from Unknown Factors for Unsupervised Video Prediction

    Vincent Le Guen, Nicolas Thome

    cs.CVarXiv:2003.01460v22020
  37. A Unifying Contrast Maximization Framework for Event Cameras, with Applications to Motion, Depth, and Optical Flow Estimation

    Guillermo Gallego, Henri Rebecq, Davide Scaramuzza

    cs.CVcs.ROarXiv:1804.01306v12018
  38. SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

    Chuan Fang, Lingteng Qiu, Yixun Liang +8

    cs.CVcs.ROarXiv:2608.27073v12026
    Summaries:한국어
  39. OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation

    Tong Wu, Jiarui Zhang, Xiao Fu +9

    cs.CVarXiv:2301.07525v22023
  40. Task-Free Continual Learning

    Rahaf Aljundi, Klaas Kelchtermans, Tinne Tuytelaars

    cs.CVcs.AIcs.LGarXiv:1812.03596v32018
  41. CLIP-ReID: Exploiting Vision-Language Model for Image Re-Identification without Concrete Text Labels

    Siyuan Li, Li Sun, Qingli Li

    cs.CVarXiv:2211.13977v42022
  42. Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-grained Image Recognition

    Heliang Zheng, Jianlong Fu, Zheng-Jun Zha +1

    cs.CVarXiv:1903.06150v22019
  43. Learning with Average Precision: Training Image Retrieval with a Listwise Loss

    Jerome Revaud, Jon Almazan, Rafael Sampaio de Rezende +1

    cs.CVarXiv:1906.07589v12019
  44. Mutual Consistency Learning for Semi-supervised Medical Image Segmentation

    Yicheng Wu, Zongyuan Ge, Donghao Zhang +4

    cs.CVcs.AIarXiv:2109.09960v42021
  45. Similarity Reasoning and Filtration for Image-Text Matching

    Haiwen Diao, Ying Zhang, Lin Ma +1

    cs.CVcs.MMarXiv:2101.01368v12021
  46. DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion

    Peize Sun, Jinkun Cao, Yi Jiang +4

    cs.CVarXiv:2111.14690v32021
  47. Explaining How a Deep Neural Network Trained with End-to-End Learning Steers a Car

    Mariusz Bojarski, Philip Yeres, Anna Choromanska +4

    cs.CVcs.LGcs.NEarXiv:1704.07911v12017
  48. Unveiling the Power of Deep Tracking

    Goutam Bhat, Joakim Johnander, Martin Danelljan +2

    cs.CVarXiv:1804.06833v12018
  49. Deep Cocktail Network: Multi-source Unsupervised Domain Adaptation with Category Shift

    Ruijia Xu, Ziliang Chen, Wangmeng Zuo +2

    cs.LGcs.CVarXiv:1803.00830v12018
  50. Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

    Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace +8

    cs.CVarXiv:2402.19479v12024
  51. Subspace Clustering by Block Diagonal Representation

    Canyi Lu, Jiashi Feng, Zhouchen Lin +2

    cs.CVarXiv:1805.09243v12018
  52. Latent Video Diffusion Models for High-Fidelity Long Video Generation

    Yingqing He, Tianyu Yang, Yong Zhang +2

    cs.CVcs.AIarXiv:2211.13221v22022
  53. HuMoR: 3D Human Motion Model for Robust Pose Estimation

    Davis Rempe, Tolga Birdal, Aaron Hertzmann +3

    cs.CVcs.LGarXiv:2105.04668v22021
  54. DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images

    Yuying Ge, Ruimao Zhang, Lingyun Wu +3

    cs.CVarXiv:1901.07973v12019
  55. Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling

    Jia Zheng, Junfei Zhang, Jing Li +3

    cs.CVarXiv:1908.00222v32019
  56. RA-UNet: A hybrid deep attention-aware network to extract liver and tumor in CT scans

    Qiangguo Jin, Zhaopeng Meng, Changming Sun +2

    cs.CVarXiv:1811.01328v12018
  57. CodeSLAM - Learning a Compact, Optimisable Representation for Dense Visual SLAM

    Michael Bloesch, Jan Czarnowski, Ronald Clark +2

    cs.CVcs.LGarXiv:1804.00874v22018
  58. MarrNet: 3D Shape Reconstruction via 2.5D Sketches

    Jiajun Wu, Yifan Wang, Tianfan Xue +3

    cs.CVcs.LGcs.NEarXiv:1711.03129v12017
  59. PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation

    Mingyang Jiang, Yiran Wu, Tianqi Zhao +2

    cs.CVarXiv:1807.00652v22018
  60. Omni-Interactive Universal Embedder

    Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui +4

    cs.AIcs.CVarXiv:2608.27044v12026