Source-linked AI summary
What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection
Parishruthi Ganesh
TL;DR
The paper asks whether richer interaction representations provide more discriminative information than coarse geometry when the learning pipeline is held fixed. It compares matched representations across weakly supervised violence-detection benchmarks and tests pre-event separability. Pose does not outperform geometry, whereas visual encoders do; substantial separation is already available before annotated onset, making video-level AUC a mixture of event and source cues.
Problem
It remains unclear whether pose’s added body-configuration detail contributes discrimination beyond coarse geometry when downstream modeling and capacity are held fixed.
Method
The study varies five interaction representations under a fixed weakly supervised early-detection pipeline, then compares visual encoders and pre-onset separation across two benchmarks.
Results
No pose representation outperforms coarse geometry; frozen appearance and context exceed geometry, while pre-onset frames retain 39–91% of above-chance separation on both benchmarks.
Takeaways & Limitations
Video-level AUC measures more than event evidence, and localizing interacting people does not improve representation performance in these evaluations.
Takeaways & Limitations
The pose comparison uses only fifteen anomalous UCF-Crime videos and leaves larger skeleton-specific architectures unevaluated.
Abstract
from arXiv · showhide
Articulated human pose provides detailed body-configuration information beyond coarse spatial relationships, but whether this detail yields greater discriminative information when the downstream pipeline is held fixed remains unclear. We examine this through early violence detection. Holding the tracker, temporal head, supervision, folds, and evaluation fixed, we compare five interaction representations spanning coarse bounding-box geometry, a matched handcrafted pose analogue, enriched pose descriptors, and a matched-capacity encoder learned from raw joints, under video-level evaluation with cluster-bootstrap intervals. No pose-based representation outperforms coarse geometry, though with fifteen anomalous videos this subset cannot rule out small effects. Extending the pipeline to frozen visual encoders, and repeating the comparison on XD-Violence (137 anomalous videos, nine times our UCF-Crime sample), person-crop appearance and whole-frame context both exceed geometry by a wide margin, yet context matches appearance on UCF-Crime and exceeds it on the larger split: cropping to the interacting people yields no advantage over encoding the whole frame. This prompts a direct test of what the benchmark measures. Scoring anomalous videos using only frames preceding the annotated onset, under a control removing sequence length as a cue, retains 39-91% of above-chance separation on both benchmarks, including for seven hand-designed geometric channels. Inspection of the tightest pre-onset windows identifies concrete provenance artifacts: editorial title cards and platform watermarks absent from the surveillance footage supplying the normal class. Video-level AUC here is thus a composite of event evidence and pre-event source cues, a shared source of discrimination that can obscure differences between representations. The diagnostic requires only annotations these benchmarks already ship.
1. Introduction
The paper isolates interaction representation quality in weakly supervised early violence detection, finding no pose advantage over coarse geometry while showing that visual context and pre-event cues strongly affect benchmark separation.
- The study asks whether pose’s extra body-configuration detail adds discrimination beyond coarse geometry when the downstream pipeline is fixed.
- Five representations are compared under one fixed tracker, temporal head, supervision, fold structure, and video-level evaluation protocol.
- No handcrafted or learned pose representation outperforms coarse bounding-box geometry in the controlled weakly supervised setting.
- The authors limit the pose conclusion to the tested descriptors, small learned encoder, and UCF-Crime pose comparison rather than interaction representations generally.
- Frozen appearance and full-frame context exceed geometry, while cropping to interacting people provides no advantage over encoding the whole frame.
- Pre-event frames retain most benchmark separation on both datasets, indicating that video-level scores combine event evidence with source-related cues.
2. Related Work
Prior interaction and pose research demonstrates expressive representations within complete recognition systems, but rarely isolates the representation’s independent contribution under matched downstream conditions.
- Existing weakly supervised anomaly and violence systems entangle interaction signals with feature encoders, temporal models, and end-to-end supervision.
- Skeleton and group-activity methods propose expressive relational encodings, but their chosen model capacities confound representation quality.
- Early event anticipation motivates the temporal setting, while focusing on prediction objectives rather than spatial representation choice.
- This paper instead varies only the interaction representation within a fixed weakly supervised early-detection pipeline.
- The broader literature’s different goals mean improvements in one setting do not establish that richer interaction representations add discriminative information under matched conditions.
3. Study Setup
The study evaluates interaction representations for weakly supervised early violence detection on an interaction-focused UCF-Crime subset and a larger XD-Violence benchmark, using video-level scoring and causal detection.
- 3.1. Task and data: UCF-Crime contributes 15 anomalous interaction videos and 150 normal videos from four predefined categories.
- 3.1. Task and data: XD-Violence adds 137 anomalous and 299 normal videos, nine times the UCF-Crime anomalous sample.
- 3.2. Detection pipeline: All representations use a shared causal temporal detection head trained with weak video-level multiple-instance supervision.
- 3.3. Evaluation: A strong full-scene detector is reported only as context because its frame-level AUC is not comparable to the study’s video-level AUCs.
- 3.3. Evaluation: The evaluation summarizes representation comparisons with video-level area under the ROC curve, while lead time and false alarms support separate failure analysis.
4. Interaction Representations
The representation ladder substitutes five pairwise interaction encodings while keeping detections, tracking, temporal modeling, training, and evaluation fixed, spanning box geometry through learned raw-joint features.
- The representation is the only experimental variable across five pairwise encodings consumed by an identical temporal detection head.
- All representations derive from shared person detections and track identities, with pose variants adding keypoints on the same tracked boxes.
- The coarse baseline uses seven normalized channels describing people count, distances, approach, alignment, overlap, and relative speed.
- The matched pose analogue replaces box quantities with skeletal counterparts so differences target skeletal signal rather than changed measurement semantics.
- Additional pose-specific channels encode strike-target proxies, elbow-angle differences, and joint acceleration, while a small raw-joint encoder tests learned retention of pair information.
- The fixed causal head uses temporal convolutions, a causal GRU, and a sigmoid classifier under identical weak supervision and fixed-length video coverage.
5. Evaluation Protocol
The study uses video-level, paired evaluation to isolate representation quality while avoiding frame-level pseudo-replication. A linear probe tests separability present in each representation, while MIL tests whether it survives the operational detection pipeline.
- Protocol: Video, rather than frame, is the unit of analysis, with whole-video cluster bootstrap uncertainty for inferential comparisons.Frame-level statistics are descriptive because correlated frames can inflate effective sample size and understate uncertainty.
- Protocol: The linear probe measures whether discriminative information is linearly present without temporal modeling, while MIL measures whether it survives the operational pipeline.The probe uses fixed-length per-video signatures and logistic regression; MIL uses the downstream detection head.
- Cross-validation: Shared repeated stratified 5-fold partitions across ten repeats pair per-video predictions across representations.The shared folds support direct paired comparisons under identical evaluation splits.
- Uncertainty: Paired ∆AUC confidence intervals resample entire videos across 2000 cluster-bootstrap resamples, and differences count as meaningful only when intervals exclude zero.The same resampled video set is scored under both representations.
- Decision criterion: Enriched pose must exceed both geometry and the tight pose analogue in the same direction under both probe and MIL to count as additional information.A single favorable fold or test is explicitly insufficient.
- Data handling: All representations use the same matched set of 165 videos, with standardized channels, zero-filled invalid frames, and an appended validity mask.For seven-channel geometry, the input contains seven features plus one validity channel.
6. Results
Across controlled interaction comparisons, pose does not outperform coarse geometry, while frozen appearance and context substantially exceed geometry. Full-frame context matches or surpasses person-crop appearance, especially on the larger benchmark, motivating scrutiny of benchmark-level cues.
- Interaction representations: 0.724 AUC was the highest linear-probe result, achieved by coarse geometry; the tight pose analogue scored 0.580.The paired P7−G7 difference was −0.153 with a 95% interval of [−0.317, −0.011], the only interval excluding zero.
- Interaction representations: 0.631 AUC from the learned raw-joint encoder remained below geometry and was reported separately without a paired probe/MIL interval.The asymmetry reflects its single end-to-end video-level evaluation rather than decomposition into the two evaluation paradigms.
- Interaction representations: Every handcrafted pose comparison against geometry was negative in central value under both probe and MIL, while pose enrichment over the tight analogue was not significant.P11−P7 was +0.047 [−0.001, 0.104] under the probe and +0.034 [−0.118, 0.173] under MIL.
- Appearance and context: +0.09 to +0.22 AUC was gained over coarse geometry by every frozen-encoder representation on both benchmarks, with intervals excluding zero in all four seeds.This was the largest and most stable effect in the study.
- Appearance and context: On UCF-Crime, APP, CTX, and their fusions were indistinguishable, whereas on XD-Violence CTX exceeded APP by 0.028–0.048 AUC under MIL and approximately 0.047 under the probe.The XD-Violence comparisons excluded zero in 4/4 seeds under both tests; the larger split contained 137 anomalous videos versus 15 in UCF-Crime.
- Appearance and context: The larger benchmark supports no advantage for cropping to interacting people over full-frame context, which was measurably better than person-crop appearance there.A single-seed UCF-Crime gain was withdrawn after reversing sign across the other seeds.
- Interpretation: A full-frame embedding matching or exceeding interaction-aware representations raises whether AUC reflects interaction-specific evidence.The paper uses this observation to motivate a direct test of what the benchmark measures.
7. What are these AUCs measuring?
The study tests whether benchmark AUC reflects event evidence by scoring pre-onset frames under a length-matched control. Substantial pre-event separation remains across representations, with artifacts demonstrating source-distribution cues.
- A full-frame encoder can separate classes without identifying interacting people, motivating a direct test for source-distribution shortcuts.The suspected cues include camera placement, venue, framing, and image quality.
- 39–84% of above-chance separation remains in seven hand-designed geometric channels before any event occurs.These channels encode only the number and arrangement of people.
- 79–91% probe retention appears on UCF-Crime and 46–70% on XD-Violence under the length-matched control.Normal videos are truncated to lengths drawn from the anomalous pre-onset distribution, removing sequence length as a class cue.
- Tight pre-onset windows contain an editorial title card and a platform watermark absent from the normal surveillance footage.These cues are class-informative without depicting violence or carrying temporal event information.
- Pre-onset separability may include legitimate escalation, but comparable retention across hand-designed channels and encoders, mostly nonviolent windows, and non-temporal artifacts indicate substantial non-event evidence.The diagnostic cannot by itself distinguish anticipation from source cues.
- Video-level AUC is a composite of event evidence and pre-event source cues, so representation differences partly measure shared, non-event discrimination.
8. Failure Analysis
The failure analysis identifies three measurable failure signatures and shows that several videos remain undetected across all representations, including frozen visual encoders.
- C1 denotes absent interaction, with two videos having zero tracked person-pairs and therefore failing the pairwise stream.
- C2 denotes atypical appearance, with three videos showing I3D centroid distances of 11.9–12.6 versus approximately 3.0 elsewhere.
- C3 denotes quiet violence, a residual failure pattern identified in the interaction-category review.The supplied passage names the pattern but truncates its further description.
- Four videos are never robustly detected across four seeds, including both C1 cases and two of three C2 cases.This includes the frozen encoders; context recovers none despite being defined on every C1 frame.
- These persistent failures are not ones that a richer pairwise descriptor would remedy.
9. Discussion
The controlled comparisons show no pose advantage over geometry, while frozen appearance and context outperform geometry without a cropping benefit. The discussion argues that benchmark AUC cannot isolate event recognition and notes important scope limits.
- 9. Discussion: Differences in the handcrafted comparison reflect representational content because model capacity and training were held constant.
- 9. Discussion: No pose representation exceeds coarse geometry; appearance and full-frame context exceed geometry, while whole-frame context exceeds person-crop appearance on XD-Violence.On UCF-Crime, appearance and context are indistinguishable from one another and from every fusion.
- 9. Discussion: Localizing interacting people yields no advantage over encoding the whole frame.
- 9. Discussion: Because 39–91% of measured separation is available before the event, video-level AUC is not event detection alone and cannot isolate pose-specific event evidence.The paper frames this as a benchmark limitation, not evidence that pose contains no interaction information.
- 9. Discussion: The pose comparison is limited to UCF-Crime’s fifteen anomalous videos and excludes high-capacity skeleton architectures by design.The small sample produces wide intervals, while larger architectures would reintroduce the controlled comparison’s capacity confound.
- 9. Discussion: Retention differs by benchmark—46–70% versus 79–91% under the probe—so its magnitude is dataset-dependent even though its presence is not.
10. Conclusion
The study finds that pose representations do not outperform coarse geometry in controlled early violence detection, while pre-event source cues complicate what video-level AUC measures.
- No pose descriptor beat coarse geometry, and frozen visual encoders improved discrimination without requiring interacting-pair localization.
- Whole-frame context exceeded person-crop appearance on the larger XD-Violence split.
- Substantial pre-onset separation remains on both benchmarks, including for handcrafted geometry, and those windows contain provenance artifacts absent from normal videos.
- Video-level AUC therefore combines event evidence with pre-event source cues, limiting conclusions about representation quality rather than proving pose is uninformative.