Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,481 to 3,540 of 18,795

  1. Robo3D: Towards Robust and Reliable 3D Perception against Corruptions

    Lingdong Kong, Youquan Liu, Xin Li +6

    cs.CVcs.ROarXiv:2303.17597v42023
  2. M2FNet: Multi-modal Fusion Network for Emotion Recognition in Conversation

    Vishal Chudasama, Purbayan Kar, Ashish Gudmalwar +3

    cs.CVcs.SDeess.ASarXiv:2206.02187v12022
  3. Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

    Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith +6

    cs.CVcs.AIcs.LGarXiv:2609.05334v12026
  4. From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

    Niu Lian, Yuting Wang, Hanshu Yao +5

    cs.CVcs.AIcs.CLarXiv:2603.01455v32026
  5. Context-aware Human Motion Prediction

    Enric Corona, Albert Pumarola, Guillem Alenyà +1

    cs.CVarXiv:1904.03419v32019
  6. RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

    Zhenxuan Fan, Bo Zhang, Yutong Lin +9

    cs.ROcs.AIcs.CVarXiv:2609.05324v12026
  7. What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

    Vivek Chavan, Pengtao Xie, Yahuan Shi +3

    cs.ROcs.AIcs.CVarXiv:2609.05376v12026
  8. Beyond Model Design: Data-Centric Training and Self-Ensemble for Gaussian Color Image Denoising

    Gengjia Chang, Xining Ge, Weijun Yuan +4

    cs.CVarXiv:2604.11468v22026
  9. Training-Free Model Ensemble for Single-Image Super-Resolution via Strong-Branch Compensation

    Gengjia Chang, Xining Ge, Weijun Yuan +4

    cs.CVarXiv:2604.11564v22026
  10. SiamMOT: Siamese Multi-Object Tracking

    Bing Shuai, Andrew Berneshawi, Xinyu Li +2

    cs.CVarXiv:2105.11595v12021
  11. OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

    Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir +8

    cs.CVarXiv:2602.17665v42026
  12. Learning Deep Bilinear Transformation for Fine-grained Image Representation

    Heliang Zheng, Jianlong Fu, Zheng-Jun Zha +1

    cs.CVarXiv:1911.03621v12019
  13. Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

    Wenjing Wang, Huan Yang, Zixi Tuo +4

    cs.CVarXiv:2305.10874v42023
  14. Invertible Denoising Network: A Light Solution for Real Noise Removal

    Yang Liu, Zhenyue Qin, Saeed Anwar +4

    eess.IVcs.CVarXiv:2104.10546v12021
  15. CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

    Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas +7

    cs.CVarXiv:2210.10716v22022
  16. S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

    Haodong Yan, Zhide Zhong, Jiaguan Zhu +10

    cs.CVcs.ROarXiv:2603.16195v22026
  17. Variational Autoencoders Pursue PCA Directions (by Accident)

    Michal Rolinek, Dominik Zietlow, Georg Martius

    cs.LGcs.CVstat.MLarXiv:1812.06775v22018
  18. FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

    Ruiteng Zhao, Wenshuo Wang, Yicheng Ma +4

    cs.ROcs.CVarXiv:2602.02142v22026
  19. Learning to drive from a world on rails

    Dian Chen, Vladlen Koltun, Philipp Krähenbühl

    cs.ROcs.CVcs.LGarXiv:2105.00636v32021
  20. Graph Degree Linkage: Agglomerative Clustering on a Directed Graph

    Wei Zhang, Xiaogang Wang, Deli Zhao +1

    cs.CVcs.SIstat.MLarXiv:1208.5092v12012
  21. Zero-Shot Visual Recognition via Bidirectional Latent Embedding

    Qian Wang, Ke Chen

    cs.CVarXiv:1607.02104v42016
  22. Interpreting Physics in Video World Models

    Sonia Joseph, Quentin Garrido, Randall Balestriero +5

    cs.CVcs.AIarXiv:2602.07050v12026
  23. A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures

    Basile Terver, Randall Balestriero, Megi Dervishi +8

    cs.CVcs.AIarXiv:2602.03604v32026
  24. RGB-D Salient Object Detection via 3D Convolutional Neural Networks

    Qian Chen, Ze Liu, Yi Zhang +3

    cs.CVarXiv:2101.10241v12021
  25. NeRF-Art: Text-Driven Neural Radiance Fields Stylization

    Can Wang, Ruixiang Jiang, Menglei Chai +3

    cs.CVcs.GRarXiv:2212.08070v12022
  26. Efficient Sharpness-aware Minimization for Improved Training of Neural Networks

    Jiawei Du, Hanshu Yan, Jiashi Feng +4

    cs.AIcs.CVcs.LGarXiv:2110.03141v22021
  27. F$^{2}$-NeRF: Fast Neural Radiance Field Training with Free Camera Trajectories

    Peng Wang, Yuan Liu, Zhaoxi Chen +5

    cs.CVcs.GRarXiv:2303.15951v12023
  28. Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly Detection

    Hui Lv, Zhongqi Yue, Qianru Sun +3

    cs.CVarXiv:2303.12369v12023
  29. Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

    Changyi Li, Yu Xiao

    cs.CVcs.AIarXiv:2609.05066v12026
  30. VPN: Learning Video-Pose Embedding for Activities of Daily Living

    Srijan Das, Saurav Sharma, Rui Dai +2

    cs.CVarXiv:2007.03056v12020
  31. Continual Semantic Segmentation via Repulsion-Attraction of Sparse and Disentangled Latent Representations

    Umberto Michieli, Pietro Zanuttigh

    cs.CVcs.AIcs.LGarXiv:2103.06342v32021
  32. LIBERO-X: Robustness Litmus for Vision-Language-Action Models

    Guodong Wang, Chenkai Zhang, Qingjie Liu +4

    cs.CVcs.AIcs.ROarXiv:2602.06556v12026
  33. Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

    Wenbin Wang, Liang Ding, Minyan Zeng +4

    cs.CVarXiv:2408.15556v12024
  34. FocalClick: Towards Practical Interactive Image Segmentation

    Xi Chen, Zhiyan Zhao, Yilei Zhang +3

    cs.CVarXiv:2204.02574v22022
  35. RAM: Recover Any 3D Human Motion in-the-Wild

    Sen Jia, Ning Zhu, Jinqin Zhong +4

    cs.CVcs.AIarXiv:2603.19929v22026
  36. The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection

    Wei Ai, Yilong Tan, Yuntao Shou +4

    cs.AIcs.CVarXiv:2601.15316v12026
  37. Rope3D: TheRoadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task

    Xiaoqing Ye, Mao Shu, Hanyu Li +5

    cs.CVarXiv:2203.13608v12022
  38. Visual News: Benchmark and Challenges in News Image Captioning

    Fuxiao Liu, Yinghan Wang, Tianlu Wang +1

    cs.CVarXiv:2010.03743v32020
  39. AdvPC: Transferable Adversarial Perturbations on 3D Point Clouds

    Abdullah Hamdi, Sara Rojas, Ali Thabet +1

    cs.CVcs.CRcs.LGarXiv:1912.00461v22019
  40. Image-to-Lidar Self-Supervised Distillation for Autonomous Driving Data

    Corentin Sautier, Gilles Puy, Spyros Gidaris +3

    cs.CVcs.LGarXiv:2203.16258v12022
  41. UNICON: Combating Label Noise Through Uniform Selection and Contrastive Learning

    Nazmul Karim, Mamshad Nayeem Rizve, Nazanin Rahnavard +2

    cs.CVcs.LGarXiv:2203.14542v42022
  42. Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

    Hulingxiao He, Zijun Geng, Yuxin Peng

    cs.CVcs.AIarXiv:2602.07605v32026
  43. B-CNN: Branch Convolutional Neural Network for Hierarchical Classification

    Xinqi Zhu, Michael Bain

    cs.CVarXiv:1709.09890v22017
  44. Learning to score the figure skating sports videos

    Chengming Xu, Yanwei Fu, Bing Zhang +3

    cs.MMcs.CVarXiv:1802.02774v32018
  45. Nighttime Dehazing with a Synthetic Benchmark

    Jing Zhang, Yang Cao, Zheng-Jun Zha +1

    cs.CVcs.LGeess.IVarXiv:2008.03864v32020
  46. Rethinking Data Augmentation for Image Super-resolution: A Comprehensive Analysis and a New Strategy

    Jaejun Yoo, Namhyuk Ahn, Kyung-Ah Sohn

    eess.IVcs.CVarXiv:2004.00448v22020
  47. Cross-Domain Few-Shot Classification via Adversarial Task Augmentation

    Haoqing Wang, Zhi-Hong Deng

    cs.CVarXiv:2104.14385v22021
  48. Simple Unsupervised Object-Centric Learning for Complex and Naturalistic Videos

    Gautam Singh, Yi-Fu Wu, Sungjin Ahn

    cs.CVcs.LGarXiv:2205.14065v12022
  49. Learning to Hash with Binary Deep Neural Network

    Thanh-Toan Do, Anh-Dzung Doan, Ngai-Man Cheung

    cs.CVarXiv:1607.05140v12016
  50. MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

    Guangheng Yang, Zhenliang Ni, Zhenkai Wu +4

    cs.CVcs.AIarXiv:2609.04947v12026
  51. Language-Conditioned World Modeling for Visual Navigation

    Yifei Dong, Fengyi Wu, Yilong Dai +10

    cs.CVcs.AIcs.ROarXiv:2603.26741v12026
  52. StableWorld: Towards Stable and Consistent Long Interactive Video Generation

    Ying Yang, Zhengyao Lv, Yujia Zeng +9

    cs.CVarXiv:2601.15281v22026
  53. One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

    Arka Pal, Rajesh Kumar, Hannes Eriksson +4

    cs.CVcs.AIcs.LGarXiv:2609.04921v12026
  54. Efficient Test-Time Adaptation of Vision-Language Models

    Adilbek Karmanov, Dayan Guan, Shijian Lu +2

    cs.CVarXiv:2403.18293v12024
  55. Sound-based Multi-Person 3D Pose Estimation

    Yusuke Oumi, Yuto Shibata, Go Irie +3

    cs.CVcs.AIcs.LGarXiv:2609.04902v12026
  56. The reliability of a deep learning model in clinical out-of-distribution MRI data: a multicohort study

    Gustav Mårtensson, Daniel Ferreira, Tobias Granberg +22

    physics.med-phcs.CVcs.LGarXiv:1911.00515v12019
  57. Unravelling Robustness of Deep Learning based Face Recognition Against Adversarial Attacks

    Gaurav Goswami, Nalini Ratha, Akshay Agarwal +2

    cs.CVarXiv:1803.00401v12018
  58. Highly Accurate Dichotomous Image Segmentation

    Xuebin Qin, Hang Dai, Xiaobin Hu +3

    cs.CVarXiv:2203.03041v42022
  59. Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

    Zhu Zhang, Zhou Zhao, Yang Zhao +3

    cs.CVarXiv:2001.06891v32020
  60. LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias

    Haian Jin, Hanwen Jiang, Hao Tan +6

    cs.CVcs.GRcs.LGarXiv:2410.17242v22024