SearcharxivSearch

arXiv subjects

Daniel Miranda

Publications and source records attributed to Daniel Miranda.

9 recordsLinked to original sources

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themselves as judges, leading to high computational costs and potential pitfalls like reward hacking and model collapse. This paper introduces a novel, model-level judge-free self-improvement framework. Our approach employs a controlled feedback mechanism while eliminating the need for MLLMs in the verification loop. We generate preference learning pairs using a controllable hallucination mechanism and optimize data quality by leveraging lightweight, contrastive language-image encoders to evaluate and reverse pairs when necessary. Evaluations across public benchmarks and our newly introduced IC dataset designed to challenge hallucination control demonstrate that our model outperforms conventional techniques. We achieve superior precision and recall with significantly lower computational demands. This method offers an efficient pathway to scalable self-improvement in MLLMs, balancing performance gains with reduced resource requirements.

cs.CL

Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

By treating visual tokens from visual encoders as text tokens, Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse visual understanding tasks, leveraging the robust architectures of Large Language Models (LLMs). However, as token counts grow, the quadratic scaling of computation in LLMs introduces a significant efficiency bottleneck, impeding further scalability. Although recent approaches have explored pruning visual tokens or employing lighter LLM architectures, the computational overhead from an increasing number of visual tokens remains a substantial challenge. In this study, we investigate the redundancy in visual computation at both the parameter and computational pattern levels within LLaVA, a representative MLLM, and introduce a suite of streamlined strategies to enhance efficiency. These include neighbor-aware visual token attention, pruning of inactive visual attention heads, and selective layer dropping for visual computations. By implementing these strategies in LLaVA, we achieve a reduction in computational demands of 88% while maintaining model performance across key benchmarks. Additionally, we validate the existence of visual computational redundancy in other MLLMs, such as Qwen2-VL-7B and InternVL-2.0-4B/8B/26B. These results present a novel pathway for MLLMs to handle dense visual tokens with minimal computational costs. Code and model checkpoints will be released to support further research.

cs.CV

Tunable narrowband Excitonic Optical Tamm States enabled by a metal-free all-organic structure

Optical Tamm States (OTS) are confined optical modes that can occur at the interface between two highly reflective structures. However, due to the strong reflectance required, their implemen-tation with highly processable and metal-free flexible materials has proven challenging. Herein, we develop the first structure supporting OTS based only on organic polymeric materials, demon-strating a photonic platform based on non-critical, widely available, and easily processable mate-rials. The structures fabricated present large areas and consist of a narrowband multi-layered polymeric Distributed Bragg Reflector (DBR) followed by a thin film of J-aggregate molecular exci-tonic material that can act as a highly reflective surface within a narrowband range. We take ad-vantage of the narrowband spectral response of the DBR and of the reflective molecular layer to tune the OTS band by varying the periodicity of the multilayer, opening the door for the fabrica-tion of OTS structures based on lightweight integrable excitonic devices with cost-effective proce-dures.

physics.optics

The limits of Near Field Immersion Microwave Microscopy evaluated by imaging bilayer graphene Moir\'{e} patterns

Molecular and atomic imaging required the development of electron and scanning probe microscopies to surpass the physical limits dictated by diffraction. Nano-infrared experiments and pico-cavity tip-enhanced Raman spectroscopy imaging later demonstrated that radiation in the visible range can surpass this limit by using scanning probe tips to access the near-field regime. Here we show that ultimate resolution can be obtained by using scanning microwave imaging microscopy to reveal structures with feature sizes down to 1~nm using a radiation of 0.1~m in wavelength. As a test material we use twisted bilayer graphene, which is not only a very important recent topic due to the discovery of correlated electron effects such as superconductivity, but also because it provides a sample where we can systematically tune a superstructure Moir\'e patterns modulation from below one up to tens of nanometers. By analyzing the tip-sample distance dynamics, we demonstrate that this ultimate 10$^8$ probe-to-pattern resolution can be achieved by using liquid immersion microscopy concepts and exquisite force control exerted on nanoscale water menisci.

cond-mat.mtrl-sci

Lattice dynamics localization in low-angle twisted bilayer graphene

A low twist angle between the two stacked crystal networks in bilayer graphene enables self-organized lattice reconstruction with the formation of a periodic domain. This superlattice modulates the vibrational and electronic structures, imposing new rules for electron-phonon coupling and the eventual observation of strong correlation and superconductivity. Direct optical images of the crystal superlattice in reconstructed twisted bilayer graphene are reported here, generated by the inelastic scattering of light in a nano-Raman spectroscope. The observation of the crystallographic structure with visible light is made possible due to lattice dynamics localization, the images resembling spectral variations caused by the presence of strain solitons and topological points. The results are rationalized by a nearly-free-phonon model and electronic calculations that highlight the relevance of solitons and topological points, particularly pronounced for structures with small twist angles. We anticipate our discovery to play a role in understanding Jahn-Teller effects and electronic Cooper pairing, among many other important phonon-related effects, and it may be useful for characterizing devices in the most prominent platform for the field of twistronics.

cond-mat.mes-hall

Path Cohomology of Locally Finite Digraphs,Hodge's Theorem and the $p$-Lazy Random Walk

The study of Markov chains on discrete spaces, such as digraphs, has captivated mathematicians in recent decades due to its interconnectedness with topology, geometry, dynamics, spectral theory, and differential equations. Furthermore, extensive exploration of these multifaceted relationships has been pursued for their practical utility in diverse fields, including machine learning and image segmentation. In recent times, these interrelations have been generalized to higher dimensions within the framework of finite-dimensional simplicial complexes. In this paper, we embark on a further extension of these concepts. Initially, we introduce a cohomology of infinite (though locally finite) digraphs in arbitrary dimensions. Subsequently, in the latter portion of this manuscript, we define a fresh family of Laplace operators and conduct an examination of their spectrum, culminating in the proof of the Hodge Decomposition Theorem within this framework. Finally, we conclude by presenting a Markov chain, the $p$-Lazy Random Walk, whose asymptotic behavior is intrinsically linked to these cohomologies, while its mixing time is related to the the spectrum of our Laplace operators. This development opens doors to numerous unexplored questions, particularly regarding potential generalizations of the Ollivier-Ricci curvature to this topology and these Laplacians.

math.PR

Invariance under quasi-isometries of subcritical and supercritical behaviour in the Boolean model of percolation

In this work we study the Poisson Boolean model of percolation in locally compact Polish metric spaces and we prove the invariance of subcritical and supercritical phases under mm-quasi-isometries. In other words, we prove that if the Poisson Boolean model of percolation is subcritical or supercritical (or exhibits phase transition) in a metric space M which is mm-quasi-isometric to a metric space N, then these phases also exist for the Poisson Boolean model of percolation in N. Then we apply these results to understand the phenomenon of phase transition in a large family of metric spaces. Indeed, we study the Poisson Boolean model of percolation in the context of Riemannian manifolds, in a large family of nilpotent Lie groups and in Cayley graphs. Also, we prove the existence of a subcritical phase in Gromov spaces with bounded growth at some scale.

math.PR

Boolean Percolation on Doubling Graphs

We consider the discrete Boolean model of percolation on graphs satisfying a doubling metric condition. We study sufficient conditions on the distribution of the radii of balls placed at the points of a Bernoulli point process for the absence of percolation, provided that the retention parameter of the underlying point process is small enough. We exhibit three families of interesting graphs where the main result of this work holds. Finally, we give sufficient conditions for ergodicity of the discrete Boolean model of percolation.

math.PR

Control Theory for Semigroups over Local Fields

Let $G$ be a 1-connected, almost-simple Lie group over a local field and $\mathcal{S}$ a subsemigroup of $G$ with non-empty interior. The action of the regular hyperbolic elements in the interior of $\mathcal{S}$ on the flag manifold $G/P$ and on the associated Euclidean building allows us to prove that the invariant control set exists and is unique. We also provide a characterization of the set of transitivity of the control sets: its elements are the fixed points of type w for a regular hyperbolic isometry, where w is an element of the Weyl group of $G$. Thus, for each w in W there is a control set $D_{w}$ and $W(\mathcal{S})$ the subgroup of the Weyl group such that the control set $D_{w}$ coincides with the invariant control set $D_{1}$ is a Weyl subgroup of $W$. We conclude by showing that the control sets are parameterized by the lateral classes $W(S)\backslash W$.

math.MG