Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
Abstract
We address the problem of training on long-tailed data forvideo action recognition. We propose to augment the training set usinga text-to-video generative model, conditioned on diverse text promptsgrounded in action profiles and training exemplars. Our approach, calledGen2Balance, converts an imbalanced training set into a balanced com-bination of real and generated video clips. To effectively learn from suchdata, we employ a two-stage training strategy that mitigates domainshift and yields significant improvements.We evaluate on long-tailed versions of standard benchmarks: UCF-101(UCF-LT) and a 100-class subset of Kinetics (K100-LT) selected to pri-oritise temporally challenging actions. Gen2Balance improves accuracyover the strongest baselines for long-tailed learning by 5.1% and 7.0% onthe respective datasets. On rare actions from the RareAct dataset (e.g.,cut keyboard ), Gen2Balance improves accuracy by 31.9%, demonstrat-ing effectiveness for scarce actions. By varying the amount of syntheticdata added, we show that partial balancing already achieves 79% of theperformance gains at 27% of the compute cost on K100-LT, highlightingthe practical scalability of Gen2Balance.