SearcharxivSearch

arXiv subjects

Guilherme Fernandes

Publications and source records attributed to Guilherme Fernandes.

3 recordsLinked to original sources

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence. This work identifies a critical structural gap in current multimodal evaluation paradigms, arguing that the reliance on Large Vision-Language Models (LVLMs) as judges is fundamentally limited by architectural biases. Our analysis reveals a profound performance dichotomy: while models may appear competent in isolated pointwise scoring, they suffer a catastrophic collapse when required to perform pairwise discrimination of temporal order. We demonstrate that this is not merely a data-scarcity issue but a structural one. Through a series of diagnostic probes, we uncover systematic positional asymmetries, specifically primacy and recency effects, where a model's judgment of a story is significantly influenced by the placement of a frame, often more than by its semantic consistency. These biases, potentially rooted in causal masking and rotary embeddings, suggest that current transformer-based judges are inherently ill-equipped for long-form visual reasoning. By exposing these blind spots, we challenge the multimedia community to move beyond snapshot-centric metrics and instead pioneer Temporally-Aware Evaluation paradigms that treat visual sequences as unified logical structures rather than unordered collections of frames.

cs.CV

Latent Beam Diffusion Models for Generating Visual Sequences

While diffusion models excel at generating high-quality images from text prompts, they struggle with visual consistency when generating image sequences. Existing methods generate each image independently, leading to disjointed narratives - a challenge further exacerbated in non-linear storytelling, where scenes must connect beyond adjacent images. We introduce a novel beam search strategy for latent space exploration, enabling conditional generation of full image sequences with beam search decoding. In contrast to earlier methods that rely on fixed latent priors, our method dynamically samples past latents to search for an optimal sequence of latent representations, ensuring coherent visual transitions. As the latent denoising space is explored, the beam search graph is pruned with a cross-attention mechanism that efficiently scores search paths, prioritizing alignment with both textual prompts and visual context. Human and automatic evaluations confirm that BeamDiffusion outperforms other baseline methods, producing full sequences with superior coherence, visual continuity, and textual alignment.

cs.CV

Against negative splitting: the case for alternative pacing strategies for elite marathon athletes in official events

Objectives: Negative splitting (i.e., finishing the race faster) is a tactic commonly employed by elite marathon athletes, even though research supporting the strategy is scarce. The presence of pacers allows the main runner to run behind a formation, preserving energy. Our aim is to show that, in the presence of pacers, the most efficient pacing strategy is positive splitting. Methods: We evaluated the performance of an elite marathon runner from an energetic standpoint, including drag values obtained through Computational Fluid Dynamics (CFD). In varying simulations for different pacing strategies, the energy for both the main runner and his pacer were conserved and the total race time was calculated. Results: In order to achieve minimum race time, the main runner must start the race faster and run behind the pacers, and when the pacers drop out, finish the race slower. Optimal race times are obtained when the protected phase is run 2.4 to 2.6% faster than the unprotected phase. Conclusion: Our results provide strong evidence that positive splitting is indeed the best pacing strategy when at least one pacer is present, causing significant time savings in official marathon events.

physics.bio-ph