Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,361 to 3,420 of 18,821

  1. Applications of Clifford's Geometric Algebra

    Eckhard Hitzer, Tohru Nitta, Yasuaki Kuroe

    math.RAcs.CVarXiv:1305.5663v12013
  2. Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

    Siyuan Liu, Xuze Zhang, Yongshun Wang +3

    cs.CVarXiv:2609.04726v12026
  3. From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction

    Xingyu Miao, Junting Dong, Qin Zhao +3

    cs.CVarXiv:2602.01661v22026
  4. Evaluating explainable artificial intelligence methods for multi-label deep learning classification tasks in remote sensing

    Ioannis Kakogeorgiou, Konstantinos Karantzalos

    cs.LGcs.CVarXiv:2104.01375v22021
  5. MogaNet: Multi-order Gated Aggregation Network

    Siyuan Li, Zedong Wang, Zicheng Liu +6

    cs.CVcs.AIarXiv:2211.03295v42022
  6. AngelFingerprint: A Traceable, Explainable, and White-Box Stealthy Watermark for Text-Guided Image Editing

    Bo-Han Kung, Futa Waseda, Ching-Chun Chang +2

    cs.CVarXiv:2609.04709v12026
  7. Understanding Neural Networks via Feature Visualization: A survey

    Anh Nguyen, Jason Yosinski, Jeff Clune

    cs.LGcs.AIcs.CVarXiv:1904.08939v12019
  8. Learning Multi-Scene Absolute Pose Regression with Transformers

    Yoli Shavit, Ron Ferens, Yosi Keller

    cs.CVarXiv:2103.11468v22021
  9. Counting Beyond Instances: A Benchmark for Group-Individual Object Counting

    Rui Wang, Junyi Huang, Jiahui Li +4

    cs.CVarXiv:2609.04716v12026
  10. Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning

    Junhao Xiao, Zhiyu Wu, Hao Lin +5

    cs.CVcs.MMarXiv:2602.21035v12026
  11. UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents

    Han Xiao, Guozhi Wang, Hao Wang +7

    cs.CVarXiv:2602.05832v12026
  12. ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features

    Jay Mahajan, Chang Liu, Rauf Makharov +3

    cs.CVarXiv:2609.04649v12026
  13. Camera Measurement of Physiological Vital Signs

    Daniel McDuff

    cs.CVcs.LGeess.IVarXiv:2111.11547v12021
  14. MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving

    Lingjun Zhang, Yujian Yuan, Changjie Wu +7

    cs.CVarXiv:2602.21952v12026
  15. Importance-Aware Low-Rank Distillation of Diffusion Transformers

    Denis Zavadski, Sebastian Heid, Damjan Kalšan +2

    cs.CVarXiv:2609.04646v12026
  16. An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

    Mahir Majid, Young Kyung Kim, Guillermo Sapiro

    cs.CVarXiv:2609.04603v12026
  17. RigNeRF: Fully Controllable Neural 3D Portraits

    ShahRukh Athar, Zexiang Xu, Kalyan Sunkavalli +2

    cs.CVarXiv:2206.06481v12022
  18. Understanding and Evaluating Racial Biases in Image Captioning

    Dora Zhao, Angelina Wang, Olga Russakovsky

    cs.CVarXiv:2106.08503v22021
  19. MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation

    Haoyuan Shi, Yunxin Li, Nanhao Deng +5

    cs.MMcs.CVarXiv:2602.23969v12026
  20. Adaptive Gated Deepfake Detection for Low-Resolution and Resource-Constrained Environments

    Vaishnavi Sen, Cody Laurie, Rashida Hasan

    cs.CVcs.LGarXiv:2609.05320v12026
  21. Diffusion Probe: Generated Image Result Prediction Using CNN Probes

    Benlei Cui, Bukun Huang, Zhizeng Ye +7

    cs.CVarXiv:2602.23783v52026
  22. Structure-Preserving Deraining with Residue Channel Prior Guidance

    Qiaosi Yi, Juncheng Li, Qinyan Dai +3

    cs.CVarXiv:2108.09079v12021
  23. DeepSteal: Advanced Model Extractions Leveraging Efficient Weight Stealing in Memories

    Adnan Siraj Rakin, Md Hafizul Islam Chowdhuryy, Fan Yao +1

    cs.CRcs.AIcs.CVarXiv:2111.04625v12021
  24. DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

    Haisheng Su, Wei Wu, Feixiang Song +3

    cs.CVarXiv:2602.13301v22026
  25. Downstream Task Inspired Underwater Image Enhancement: A Perception-Aware Study from Dataset Construction to Network Design

    Bosen Lin, Feng Gao, Yanwei Yu +2

    cs.CVeess.IVarXiv:2603.01767v12026
  26. Pneumonia Detection on chest X-ray images Using Ensemble of Deep Convolutional Neural Networks

    Alhassan Mabrouk, Rebeca P. Díaz Redondo, Abdelghani Dahou +2

    eess.IVcs.CVcs.LGarXiv:2312.07965v12023
  27. DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided Network

    Yeying Jin, Aashish Sharma, Robby T. Tan

    cs.CVarXiv:2207.10434v22022
  28. Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning

    Qi Li, Xinchao Wang

    cs.CVarXiv:2601.17566v12026
  29. SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis

    Yuqing Yang, Alexander Schmatz, Zhaozhao Ma +3

    cs.CVcs.LGarXiv:2609.05174v12026
  30. SAM-Assisted Remote Sensing Imagery Semantic Segmentation with Object and Boundary Constraints

    Xianping Ma, Qianqian Wu, Xingyu Zhao +3

    cs.CVarXiv:2312.02464v22023
  31. Total Variation Regularized Tensor RPCA for Background Subtraction from Compressive Measurements

    Wenfei Cao, Yao Wang, Jian Sun +4

    cs.CVarXiv:1503.01868v42015
  32. LookThere! Sparse Vision by Reinforced Selection

    Sreehari Rammohan, Yousef Yassin, Anthony Fuller +3

    cs.CVcs.LGarXiv:2609.04698v12026
  33. Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective

    Qiyan Zhao, Xiaofeng Zhang, Shuochen Chang +7

    cs.CVarXiv:2601.20520v12026
  34. Memory-Efficient Implementation of DenseNets

    Geoff Pleiss, Danlu Chen, Gao Huang +3

    cs.CVarXiv:1707.06990v12017
  35. Lightweight image super-resolution with enhanced CNN

    Chunwei Tian, Ruibin Zhuge, Zhihao Wu +4

    eess.IVcs.CVarXiv:2007.04344v32020
  36. Compressive Hyperspectral Imaging with Side Information

    Xin Yuan, Tsung-Han Tsai, Ruoyu Zhu +3

    cs.CVarXiv:1502.06260v12015
  37. Anomaly Detection in Road Traffic Using Visual Surveillance: A Survey

    Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy

    cs.CVarXiv:1901.08292v12019
  38. An Alternative Probabilistic Interpretation of the Huber Loss

    Gregory P. Meyer

    stat.MLcs.CVcs.LGarXiv:1911.02088v32019
  39. LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

    Zekun Li, Sizhe An, Chengcheng Tang +7

    cs.CVarXiv:2602.12370v22026
  40. Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization

    Xingjian Diao, Zheyuan Liu, Chunhui Zhang +6

    cs.CVcs.CLarXiv:2601.04442v22026
  41. On the Generalization Capacities of MLLMs for Spatial Intelligence

    Gongjie Zhang, Wenhao Li, Quanhao Qian +4

    cs.CVcs.LGarXiv:2603.06704v12026
  42. IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

    Yinghao Tang, Xueding Liu, Boyuan Zhang +13

    cs.LGcs.CVarXiv:2601.04498v22026
  43. LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

    Yichen Zhu, Minjie Zhu, Ning Liu +3

    cs.CVcs.CLarXiv:2401.02330v42024
  44. Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled Hardware

    Aayush Marasini, Zhaoxian Zhou

    cs.ARcs.CVcs.LGarXiv:2609.04705v12026
  45. ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

    Javier del Pino, Salvador Rodríguez, Alejandro Garabito +2

    cs.CVcs.AIarXiv:2609.03756v12026
  46. Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks

    Le Hou, Chen-Ping Yu, Dimitris Samaras

    cs.CVarXiv:1611.05916v42016
  47. Knowledge distillation from multi-modal to mono-modal segmentation networks

    Minhao Hu, Matthis Maillard, Ya Zhang +4

    cs.CVcs.AIstat.MLarXiv:2106.09564v12021
  48. Unpaired Brain MR-to-CT Synthesis using a Structure-Constrained CycleGAN

    Heran Yang, Jian Sun, Aaron Carass +4

    cs.CVarXiv:1809.04536v12018
  49. Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

    Haoyuan Li, Qihang Cao, Tao Tang +6

    cs.CVarXiv:2602.06037v52026
  50. Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion

    Yang Yang, Tianyi Zhang, Wei Huang +6

    cs.CVeess.IVarXiv:2603.13405v12026
  51. Towards Automated Deep Learning: Efficient Joint Neural Architecture and Hyperparameter Search

    Arber Zela, Aaron Klein, Stefan Falkner +1

    cs.LGcs.AIcs.CVarXiv:1807.06906v12018
  52. Hidden In Plain Gaze: Gaze Representations as Privacy Controls for Utility and Re-identification Risk in XR

    Cory Ilo, Brendan-David John, Doug A. Bowman

    cs.CVcs.ETcs.HCarXiv:2609.04592v12026
  53. Efficient Chest X-ray Representation Learning via Semantic-Partitioned Contrastive Learning

    Wangyu Feng, Shawn Young, Lijian Xu

    cs.CVarXiv:2603.07113v22026
  54. Adversarial attacks and defenses in explainable artificial intelligence: A survey

    Hubert Baniecki, Przemyslaw Biecek

    cs.CRcs.AIcs.CVarXiv:2306.06123v42023
  55. Data-driven Flood Emulation: Speeding up Urban Flood Predictions by Deep Convolutional Neural Networks

    Zifeng Guo, Joao P. Leitao, Nuno E. Simoes +1

    cs.CVcs.CYcs.LGarXiv:2004.08340v22020
  56. StreamReady: Learning What to Answer and When in Long Streaming Videos

    Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat

    cs.CVarXiv:2603.08620v12026
  57. Learning Transferrable Knowledge for Semantic Segmentation with Deep Convolutional Neural Network

    Seunghoon Hong, Junhyuk Oh, Bohyung Han +1

    cs.CVarXiv:1512.07928v12015
  58. Interpretable Convolutional Neural Networks via Feedforward Design

    C. -C. Jay Kuo, Min Zhang, Siyang Li +2

    cs.CVarXiv:1810.02786v22018
  59. High-Fidelity Image Generation With Fewer Labels

    Mario Lucic, Michael Tschannen, Marvin Ritter +3

    cs.LGcs.CVstat.MLarXiv:1903.02271v22019
  60. Automatic Brain Tumor Segmentation using Convolutional Neural Networks with Test-Time Augmentation

    Guotai Wang, Wenqi Li, Sebastien Ourselin +1

    cs.CVarXiv:1810.07884v22018