Search arXiv⌕ Search

arXiv subjects

Mahsa Mohammadi

Publications and source records attributed to Mahsa Mohammadi.

2 recordsLinked to original sources

Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges

Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different stages, but reported gains are difficult to interpret when task formulation, visual pretraining, supervision, decoder design, and evaluation code change simultaneously. The central contribution of this review is an evidence-aware design map that crosses task regime with the point at which language-derived information intervenes. We characterise task regimes along six axes. These axes organise the literature into five broad task families: single-action, sequence, object-interaction, cross-view, and planning-oriented settings. C1-C3 locate interventions in context construction, goal/intention modelling, and future decoding, while C4 is treated as an adjacent, emerging grounding/executability extension. Unlike a generic processing pipeline, the map links each intervention to an appropriate counterfactual, failure diagnosis, and permissible evidence claim. Supporting contributions include a protocol-level audit of Ego4D-LTA and EPIC-KITCHENS-100, a multidimensional evidence profile, and the Backbone-Aware Comparison and Ablation Protocol (BCAP). The unresolved EK-100 record is treated as a reporting-comparability case study and is not used as a leaderboard. Evidence for LLM benefits, goal ambiguity, and horizon effects is therefore formulated as testable hypotheses requiring matched validation, not as causal conclusions. The accompanying package contains the coded evidence, source locators, protocol metadata, and versioned catalogue used in the review.

cs.CV↗

Reliable Egocentric Action Anticipation via Temporal Reliability Suppression and Compositional Graph Decoding

Wearable action anticipation systems must remain reliable despite missing frames, masking, and sensor noise, yet existing egocentric anticipation methods largely assume clean observations. We identify two complementary failure modes under temporal corruption: unreliable temporal evidence during encoding and implausible, low-support verb-noun compositions during decoding. We address them with a lightweight framework combining Temporal Reliability Suppression (TRS) and Robust Verb-Noun Graph (RVG) decoding. TRS predicts a per-frame suppression score from the projected input embedding and uses it as a learned key-side attention penalty at every encoder block and to derive reliability-weighted temporal pooling. RVG re-ranks verb-noun pairs using a PMI-based compatibility graph constructed from training labels. Under corruption-augmented training, TRS+CA+RVG reaches 29.1% average corrupted accuracy and 88.2% relative robustness across six corruptions, including three mechanisms absent during training, while reducing rare verb-noun predictions from 15.4% to 1.1%. Multi-seed and diagnostic experiments show that TRS responds to synthetic masking; shuffled-graph and frequency-only controls further indicate that RVG gains depend on genuine pairwise compatibility rather than marginal-frequency effects alone.

cs.CV↗