SearcharxivSearch

arXiv subjects

Clark Peng

Publications and source records attributed to Clark Peng.

3 recordsLinked to original sources

DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation

Contact languages like English exhibit rich regional variations in the form of dialects, which are often used by dialect speakers interacting with generative models. However, can multimodal generative models effectively produce content given dialectal textual input? In this work, we study this question by constructing a new large-scale benchmark spanning six common English dialects. We work with dialect speakers to collect and verify over 4200 unique prompts and evaluate on 17 image and video generative models. Our automatic and human evaluation results show that current state-of-the-art multimodal generative models exhibit 32.26% to 48.17% performance degradation when a single dialect word is used in the prompt. Common mitigation methods such as fine-tuning and prompt rewriting can only improve dialect performance by small margins (< 7%), while potentially incurring significant performance degradation in Standard American English (SAE). To this end, we design a general encoder-based mitigation strategy for multimodal generative models. Our method teaches the model to recognize new dialect features while preserving SAE performance. Experiments on models such as Stable Diffusion 1.5 show that our method is able to simultaneously raise performance on five dialects to be on par with SAE (+34.4%), while incurring near zero cost to SAE performance.

cs.CL

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world actions remains unclear (e.g., playing tennis, backflip). Existing benchmarks suffer from limitations such as limited size, lack of human evaluation, sim-to-real gaps, and absence of fine-grained physical rule analysis. To address this, we introduce VideoPhy-2, an action-centric dataset for evaluating physical commonsense in generated videos. We curate 200 diverse actions and detailed prompts for video synthesis from modern generative models. We perform human evaluation that assesses semantic adherence, physical commonsense, and grounding of physical rules in the generated videos. Our findings reveal major shortcomings, with even the best model achieving only 22% joint performance (i.e., high semantic and physical commonsense adherence) on the hard subset of VideoPhy-2. We find that the models particularly struggle with conservation laws like mass and momentum. Finally, we also train VideoPhy-AutoEval, an automatic evaluator for fast, reliable assessment on our dataset. Overall, VideoPhy-2 serves as a rigorous benchmark, exposing critical gaps in video generative models and guiding future research in physically-grounded video generation. The data and code is available at https://videophy2.github.io/.

cs.CV

Boundary Density Likelihood for Direct Event-Time Supervision

Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models are trained for samplewise segmentation and only convert predicted states into events after training. We ask whether training directly for the evaluated output improves detection. Boundary Density Likelihood (BDL) assigns one unit of target mass to each annotated event, preserves that mass through smoothing and temporal downsampling, and uses a Poisson objective to estimate expected event mass in each output bin; local peaks become ranked detections. In a prespecified five-fold nested sleep study, BDL-Hard raises pooled out-of-fold mAP from 0.586 to 0.705 over interval segmentation (+11.9 percentage points; 95% interval [10.8, 13.0]) and strict one-minute AP from 0.071 to 0.286 (4.0x; +21.5 points), improving on every outer fold. A separate held-out rerun reproduces the direction of the effect. A matched boundary-BCE detector reaches 0.702 mAP, showing that direct boundary supervision and event decoding account for most of the gain, with a smaller contribution from the Poisson objective. The same trend appears with an offline convolutional model; causal, Transformer, and patient-grouped seizure experiments leave broader generalization unresolved. Our results support training on event times when timestamps define the evaluated output.

cs.AI