Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,561 to 13,620 of 18,776

  1. Generative Multimodal Models are In-Context Learners

    Quan Sun, Yufeng Cui, Xiaosong Zhang +8

    cs.CVarXiv:2312.13286v22023
  2. Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases

    Christopher Clark, Mark Yatskar, Luke Zettlemoyer

    cs.CLcs.CVcs.LGarXiv:1909.03683v12019
  3. PoseTrack: A Benchmark for Human Pose Estimation and Tracking

    Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov +4

    cs.CVarXiv:1710.10000v22017
  4. InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

    Jiale Xu, Weihao Cheng, Yiming Gao +3

    cs.CVarXiv:2404.07191v22024
  5. DeeperGCN: All You Need to Train Deeper GCNs

    Guohao Li, Chenxin Xiong, Ali Thabet +1

    cs.LGcs.CVstat.MLarXiv:2006.07739v12020
  6. LangSplat: 3D Language Gaussian Splatting

    Minghan Qin, Wanhua Li, Jiawei Zhou +2

    cs.CVarXiv:2312.16084v22023
  7. PARE: Part Attention Regressor for 3D Human Body Estimation

    Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges +1

    cs.CVarXiv:2104.08527v22021
  8. Pixel Difference Networks for Efficient Edge Detection

    Zhuo Su, Wenzhe Liu, Zitong Yu +5

    cs.CVarXiv:2108.07009v12021
  9. DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image Classification

    Hongrun Zhang, Yanda Meng, Yitian Zhao +4

    cs.CVcs.AIcs.LGarXiv:2203.12081v12022
  10. Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering

    Haoyuan Gao, Junhua Mao, Jie Zhou +3

    cs.CVcs.CLcs.LGarXiv:1505.05612v32015
  11. Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image

    Denis Tome, Chris Russell, Lourdes Agapito

    cs.CVarXiv:1701.00295v42017
  12. Face Aging With Conditional Generative Adversarial Networks

    Grigory Antipov, Moez Baccouche, Jean-Luc Dugelay

    cs.CVarXiv:1702.01983v22017
  13. See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks

    Xiankai Lu, Wenguan Wang, Chao Ma +3

    cs.CVarXiv:2001.06810v12020
  14. Online Continual Learning in Image Classification: An Empirical Survey

    Zheda Mai, Ruiwen Li, Jihwan Jeong +3

    cs.LGcs.CVarXiv:2101.10423v42021
  15. FOTS: Fast Oriented Text Spotting with a Unified Network

    Xuebo Liu, Ding Liang, Shi Yan +3

    cs.CVarXiv:1801.01671v22018
  16. Understanding Diffusion Models: A Unified Perspective

    Calvin Luo

    cs.LGcs.CVarXiv:2208.11970v12022
  17. DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

    Lu Ling, Yichen Sheng, Zhi Tu +17

    cs.CVcs.AIarXiv:2312.16256v22023
  18. CVE-SAI: Counterfactual Visual Evidence-Guided Selective Attribute Indexing for Risk-Controlled E-commerce Search

    Xiaolong Sun, Qichao Wang, Hangyu Li +1

    cs.AIcs.CVarXiv:2608.25023v12026
  19. Stochastic Separability of Embedding Manifolds

    Liqing Zhang

    cs.LGcs.CVarXiv:2608.22874v12026
  20. Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges

    Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar

    cs.CVarXiv:2608.23845v12026
  21. Investigating Relational Reasoning in VLMs

    Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan +1

    cs.CVarXiv:2608.23518v12026
  22. MorphoCLIP: Text-Supervised Contrastive Learning for Perturbation Matching in Cell Painting Images

    Sukhrobbek Ilyosbekov, Shubham Gajjar, Rongfei Jin

    cs.CVarXiv:2608.22690v12026
  23. Targeting the Attention Heads Behind Object Hallucination in LLaVA

    Armaan Sandhu, Abhilasha Senapati, Hima Kammachi

    cs.CVcs.AIarXiv:2608.24966v12026
  24. FlashNormal: Detailed Surface Normal Estimation from Flash and No-Flash Images

    Ruiyang Chen, Feiran Li, Heng Guo +1

    cs.CVarXiv:2608.25360v12026
  25. Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

    Jheng-Ling Lee, Shang-Tse Chen

    cs.CVcs.LGarXiv:2608.24093v12026
  26. When Should a Network Emit Geometry, and When Should It Detect It? Readout, Reconciliation, and Representation in Floorplan Vectorization

    He Zhang

    cs.CVarXiv:2608.25608v12026
  27. PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction

    Mary Raymond

    cs.CVcs.AIarXiv:2608.25401v12026
  28. RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

    Bojia Zi, Xiaoyan Yang, Yu Zhou +7

    cs.CVarXiv:2608.26101v12026
  29. A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

    Kaichen Li, Zhilin Zhu, Jianhao Huang +7

    cs.CVcs.AIarXiv:2608.26095v12026
  30. DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

    Tuo Chen, Jie Gui, Minjing Dong +4

    cs.CVarXiv:2608.25851v12026
  31. Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution

    Ren Wang, Yung-Yu Chuang

    cs.CVarXiv:2608.25998v12026
  32. 4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

    Yueen Ma, Zenglin Xu, Irwin King

    cs.CVarXiv:2608.25956v12026
  33. TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

    Yuqiang Lin, Yan Shi, Sam Lockyer +5

    cs.CVcs.AIarXiv:2608.25935v12026
  34. UltraPIPS: Improving model perception in B-mode ultrasound with foundation models

    Tal Grutman, Tali Ilovitsh

    cs.CVeess.IVarXiv:2608.26033v12026
  35. Visual General Intelligence: A White Paper

    Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18

    cs.CVarXiv:2608.25924v12026
  36. Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

    Juan Iñaki Larrea, Lucas Mansilla, Enzo Ferrante

    cs.CVarXiv:2608.25810v12026
  37. Embedding NDRE Trajectories into Contrastive Learning for Label-Free, Physiology-Aware Crop-Stress Staging and DSS Outputs

    Shafqaat Ahmad

    cs.CVarXiv:2608.25888v12026
  38. LUTSeg: A Longitudinal Multi-Expert Dataset for Ulcer Tissue Segmentation

    Karen Sanchez, Carlos Hinojosa, Albert A. Ávila +5

    cs.CVarXiv:2608.25866v12026
  39. Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

    Jiaming Zhou, Qihang Zhang, Gangwei Xu +9

    cs.ROcs.CVarXiv:2608.26103v12026
  40. Auditable CT Phenotyping Through Report-derived Radiological Observations

    Riga Wu, Walter Witschey, Yicheng Li +9

    cs.CVarXiv:2608.25948v12026
  41. Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future

    Grace W. Lindsay

    q-bio.NCcs.CVcs.NEarXiv:2001.07092v22020
  42. Learning Late, Guiding Early: Timestep-Decoupled Semantic Guidance for Fair Face Generation

    Subir Kumar Parida, Rajbabu Velmurugan, Ketan Kotwal +2

    cs.CVarXiv:2608.25862v12026
  43. Steer the Sampling, Not the Kernel Grid: Geometry-Guided Sampling Operator for Volumetric Segmentation

    Sizhe Wang, Himashi Peiris, Zhaolin Chen

    cs.CVarXiv:2608.25819v12026
  44. TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection

    Qiangqiang Zhou, Jiacong Yu, Jiawei Xu +3

    cs.CVarXiv:2608.25808v12026
  45. THA-Flow Generative Model: Prosthesis Geometry Prediction from Preoperative CT

    Yiping Wang, Jie Li, Jingyu Shen +1

    cs.CVarXiv:2608.25845v12026
  46. FRAME: separating sampling variation from representational cause in medical imaging fairness

    Mahshad Lotfinia, Daniel Truhn, Andreas Maier +1

    cs.CVcs.AIcs.LGarXiv:2608.25981v12026
  47. Multi-Oriented Text Detection with Fully Convolutional Networks

    Zheng Zhang, Chengquan Zhang, Wei Shen +3

    cs.CVarXiv:1604.04018v22016
  48. InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control

    Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli +5

    cs.CVarXiv:2608.25734v12026
  49. Deep Learning Segmentation of Diffusion-Weighted MRI Acute Ischaemic Stroke: A Pragmatic Evaluation Across Three Datasets

    Atle Bjørnerud, Till Schellhorn, Thor H. Skattør +4

    cs.CVarXiv:2608.25675v12026
  50. CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery

    Yuanpei Liu, Zhenqi He, Jialu Tang +1

    cs.CVarXiv:2608.25692v12026
  51. Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Fuxiao Liu, Kevin Lin, Linjie Li +3

    cs.CVcs.AIcs.CEarXiv:2306.14565v42023
  52. Precipitation Downscaling Using Foundation Model-Conditioned Diffusion

    Victor Nascimento Ribeiro, Jorge Guevara, Jorge Sebastian Moraga +7

    cs.CVcs.LGphysics.ao-pharXiv:2608.25858v12026
  53. State of the Art on Neural Rendering

    Ayush Tewari, Ohad Fried, Justus Thies +16

    cs.CVcs.GRarXiv:2004.03805v12020
  54. MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

    Hao Yin, Paritosh Parmar, Lijun Gu +6

    cs.CVcs.AIcs.ETarXiv:2608.26094v12026
  55. AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis

    Yudong Guo, Keyu Chen, Sen Liang +3

    cs.CVarXiv:2103.11078v32021
  56. MIMONet: Multi-scale Input and Multi-scale Output Network for Salient Object Detection

    Zhaojian Yao, Wei Gao, Tiesong Zhao +2

    cs.CVarXiv:2608.25733v12026
  57. LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

    Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase +3

    cs.CVarXiv:2608.25729v12026
  58. Difficulty-Aware Sample Allocation for Adaptive Data Augmentation in Semantic Segmentation

    Olasimbo Ayodeji Arigbabu, Abimbola Ismail Arigbabu

    cs.CVcs.AIarXiv:2608.25710v12026
  59. VITAL: VIsual Tracking via Adversarial Learning

    Yibing Song, Chao Ma, Xiaohe Wu +6

    cs.CVarXiv:1804.04273v12018
  60. Controlling for Omitted Variable Bias in Deep Neural Networks

    Manuel Pfeuffer, Roshan Prakash Rane, Kerstin Ritter +1

    stat.MEcs.CVcs.LGarXiv:2608.25930v12026