
Apple Machine Learning Research describes evaluating the quality of video captions using multiple-choice question answers. The approach is presented in the paper “Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering.”
According to the publisher's synopsis, common metrics compare generated text to reference descriptions. Apple notes that this principle poorly accounts for cases where a single video permits multiple correct descriptions and typically yields only a one-dimensional quality score.
editorial commentary
Why it matters
A probable consequence is continued attention to evaluating video captions through content verification, not only word matching. The next observable signal will be the publication of details of the methodology, experiments, and comparisons with existing metrics. Substantial uncertainty remains: only a synopsis is provided, without results and a full description of the approach.