SearcharxivSearch

arXiv subjects

Yuhao Zhao

Publications and source records attributed to Yuhao Zhao.

At least 19 recordsLinked to original sources

Supersaturation for Eventown via Generator Switching

An eventown family is a family of even-sized subsets of $[n]$ in which every two distinct members have an even-sized intersection. A classical theorem of Berlekamp and Graver shows that the maximum size of such a family is $2^{\lfloor n/2\rfloor}$. The supersaturation problem for eventown asks how many odd-intersection pairs must occur when this extremal bound is exceeded. For a family $\mathcal F$ of even-sized subsets of $[n]$, let $e(\mathcal F)$ denote the number of unordered pairs whose intersection size is odd. O'Neill conjectured that if $|\mathcal F|=2^{\lfloor n/2\rfloor}+s$, then $e(\mathcal F)\ge s\,2^{\lfloor n/2\rfloor-1}$ for \[ 1\le s\le 2^{\lfloor n/2\rfloor}-2^{\lfloor n/4\rfloor}. \] Previously, the conjecture was known for $s=1,2$, and, for $s\le 2^{\lfloor n/8\rfloor}/n$ with $n$ sufficiently large. We prove the conjectured bound for \[ 1\le s\le \frac{2^{\lfloor n/2\rfloor}}{26}, \] extending the known range to a fixed positive proportion of the extremal eventown size. The bound is sharp throughout this range. As further consequences, we derive a lower bound valid for arbitrary excess $s$, which improves the previously known estimate in an additional range. We also establish stability and removal results for families of extremal size satisfying $e(\mathcal F)<2^{\lfloor n/2\rfloor-1}$, showing that such a family is close to an extremal eventown family and can be made eventown by deleting a small number of its members.

math.CO

Bounds and constructions for constant/low-power error-correcting cooling codes

The low-power error-correcting cooling (LPECC) codes and constant-power error-correcting cooling (CPECC) codes, introduced in [IEEE Trans. Inf. Theory, 64 (2018), 3062--3085; 66 (2020), 4804--4818], respectively, are two coding schemes designed to simultaneously control the peak temperature and average power consumption of on-chip buses while providing error-correction capability for transmitted information. This paper establishes new upper bounds for both $(n,1,w,w-2)$-CPECC codes and $(n,t,w,w-2)$-CPECC codes using graph-theoretic techniques, and constructs several new families of optimal CPECC codes using combinatorial configurations. Moreover, it completely resolves the conjecture concerning CPECC codes posed in [IEEE Trans. Inf. Theory, DOI: 10.1109/TIT.2026.3721101]. Finally, we derive a new upper bound for $(n,t,w,w-2)$-LPECC codes by probabilistic method, along with new optimal families, and establish the relationship between optimal $(n,t,w,w-2)$-LPECC codes and optimal $(n+1,t,w,w-2)$-CPECC codes.

cs.IT

A superlogarithmic saving for Oddtown modulo composite numbers

Let $f_{\ell}(n)$ be the largest size of a family $\mathcal{A}\subseteq2^{[n]}$ such that no member has size divisible by $\ell$, while the intersection of every two distinct members has size divisible by $\ell$, and let $\omega(\ell)$ denote the number of distinct prime divisors of $\ell$. For any prime power $\ell$, the classical answer is $f_{\ell}(n)=n$. When $\omega(\ell)\geq 2$, Bukh, Chao, and Zheng recently proved $\omega(\ell)n-O_{\ell}\left(n^{\frac{\omega(\ell)-2}{\omega(\ell)-1}}(\log n)^{C_{\ell}}\right)\leq f_{\ell}(n)\leq\omega(\ell)n-2\omega(\ell)\log n+11$ for some $C_{\ell}>0$. When $\ell$ has at least two distinct odd prime divisors, they further used Fourier analysis to improve the upper bound to $f_{\ell}(n)\leq\omega(\ell)n-(2\omega(\ell)+\varepsilon_{\ell})\log n$ for some $\varepsilon_{\ell}>0$, provided that $n$ is sufficiently large in terms of $\ell$ . For every fixed $\ell$ with $\omega(\ell)\geq2$, we prove \[ f_{\ell}(n)\leq\omega(\ell)n-\Omega_{\ell}(\log n\log\log n) \] for large $n$. The upper bound relies on a submatrix lemma of Bhowmick, Dvir, and Lovett, which is based on the bounded-torsion polynomial Freiman--Ruzsa conjecture recently proved by Gowers, Green, Manners, and Tao.

math.CO

Uni-XAS: Alignment-Driven Bidirectional Multimodal Learning for X-ray Absorption Spectroscopy

X-ray absorption spectroscopy (XAS) is a key technique for probing local atomic environments, yet learning based modeling must bridge two heterogeneous modalities: 1D continuous spectra and 3D atomic structures. Existing approaches typically decouple forward spectrum prediction and inverse structure inference into separate regression tasks, hindering shared representation learning. Moreover, severe permutation ambiguity among identical atoms often limits inverse modeling to coarse structure descriptors rather than explicit 3D structure generation. In this work, we present Uni-XAS, a unified benchmark and learning framework that reframes bidirectional XAS modeling as a cross-modal alignment and conditional generation problem. We first propose XASLip, an alignment recipe coupling a physics-aware spectral encoder with an absorberaware manifold optimization strategy to resolve fine-grained intra-element coordination variations. Building upon this shared latent space, we formulate forward prediction as anchored absolute-spectrum generation via retrieval-augmented decoding, effectively preventing physical scale collapse and energy drift. For the inherently ill-posed inverse problem, we introduce Permutation-Rectified Flow Matching, which integrates type-wise optimal transport into a continuous generative flow to provide a principled solution to ligand permutation ambiguity without relying on heavy high-order equivariant architectures. Evaluated on a largescale standardized benchmark of 328,839 structure-spectrum pairs, Uni-XAS demonstrates strong performance in cross-modal retrieval, accurate absolute-spectrum prediction, and composition-conditional 3D structure generation, establishing a scalable, reproducible, and protocol-consistent foundation for multimodal learning and standardized evaluation in scientific spectroscopy.

cond-mat.mtrl-sci

Lower bounds on the strength of the determinant

We establish new lower bounds for the strength and partition rank of the determinant. For every prime $p$, we prove the exact identity \[ \operatorname{str}(\mathrm{det}_p)=p. \] A weak monotonicity argument, combined with a bound for gaps between consecutive primes, then gives $\operatorname{str}(\mathrm{det}_n)\ge (1-o(1))n^{0.475}$ for sufficiently large $n$. Since the Birch rank of $\mathrm{det}_n$ is always $4$, this gives the first explicit family showing that the dependence on the degree in bounds for strength in terms of Birch rank is unavoidable. Viewing $\mathrm{det}_n$ as an $n$-linear form in its columns, we also prove that its partition rank is at least the largest prime not exceeding $n$. Consequently, \[ n-n^{0.525}\le \operatorname{prk}(\mathrm{det}_n)\le n \] for all sufficiently large $n$, and hence the partition rank of the determinant is $n-o(n)$. The proof introduces an intersection-theoretic method for lower-bounding strength: a short strength decomposition produces a nowhere-vanishing section of a split vector bundle on the complement of the determinantal hypersurface, while a nonzero top Chern class in the Chow ring of $\mathrm{PGL}_n$ obstructs such a section.

math.AC

Solving the inverse problem of X-ray absorption spectroscopy via physics-informed deep learning

Resolving transient atomic configurations in non-crystalline or dynamic environments remains a fundamental bottleneck in the physical sciences. While X-ray absorption spectroscopy (XAS) is a premier probe of local structure, inverting spectra into structural descriptors is a notoriously ill-posed problem due to inherent many-to-one mapping. Here, we present the Spectral Pattern Translator (SPT), a physics-informed deep learning framework that establishes a robust bridge between large-scale theoretical datasets and experimental reality. Our strategy exploits the Fourier duality between spectral energy oscillations and spatial scattering paths to overcome the "simulation-to-experiment" gap. By decomposing spectra into frequency domains, SPT effectively isolates robust structural coordination signals from the destabilizing noise inherent in experimental data. Trained on a massive library of diverse atomic environments, this approach achieves state-of-the-art accuracy in resolving continuous phase transitions in battery cathodes and deciphering local order in amorphous materials. With millisecond-scale latency, SPT removes the primary computational barrier to autonomous materials discovery, establishing a robust, noise-resilient engine for closed-loop robotic chemistry.

cond-mat.mtrl-sci

MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding

Point cloud completion aims to recover complete 3D geometry from partial observations caused by limited viewpoints and occlusions. Existing learning-based works, including 3D Convolutional Neural Network (CNN)-based, point-based, and Transformer-based methods, have achieved strong performance on synthetic benchmarks. However, due to the limitations of modality, scalability, and generative capacity, their generalization to novel objects and real-world scenarios remains challenging. In this paper, we propose MGPC, a generalizable multimodal point cloud completion framework that integrates point clouds, RGB images, and text within a unified architecture. MGPC introduces an innovative modality dropout strategy, a Transformer-based fusion module, and a novel progressive generator to improve robustness, scalability, and geometric modeling capability. We further develop an automatic data generation pipeline and construct MGPC-1M, a large-scale benchmark with over 1,000 categories and one million training pairs. Extensive experiments on MGPC-1M and in-the-wild data demonstrate that the proposed method consistently outperforms prior baselines and exhibits strong generalization under real-world conditions.

cs.CV

Observing Laughlin's pump using quantized edge states in graphene

Laughlin's thought experiment of quantized charge pumping is central to understanding the integer quantum Hall effect (IQHE) and the topological origin of its conductance quantization. Its direct experimental observation, however, has been hindered by the difficulty of realizing clean electronic edges. We address this by fabricating ultra-small, lithographically defined contacts on graphene. This creates a Corbino-equivalent system, with well-confined inner edge states. Crucially, the small contact size induces strong energy quantization of the edge states. This quantization allows us to directly resolve the spectral flow associated with Laughlin's pump. By tracing the finite-size resonances of the inner edge, we observe clear oscillations in conductance as a function of magnetic field and carrier density. The oscillation period scales with contact size, consistent with quantized charge transfer. Thus, our results provide a direct observation of the spectral flow underlying Laughlin's pump. The simplicity of the graphene platform makes this approach scalable and robust for exploring fundamental topological effects.

cond-mat.mes-hall

Counting cliques with prescribed intersection sizes

We study the generalized Tur\'an problem regarding cliques with restricted intersections, which highlights the motivation from extremal set theory. Let $L=\{\ell_1,\dots,\ell_s\}\subset [0,r-1]$ be a fixed integer set with $|L|\notin \{1,r\}$ and $\ell_1<\dots<\ell_s$, and let $\Psi_r(n,L)$ denote the maximum number of $r$-cliques in an $n$-vertex graph whose $r$-cliques are $L$-intersecting as a family of $r$-subsets. Helliar and Liu recently initiated the systematic study of the function $\Psi_r(n,L)$ and showed that $\Psi_r(n,L)\le \left(1-\frac{1}{3r}\right) \prod_{\ell\in L}\frac{n-\ell}{r-\ell}$ for large $n$, improving the trivial bound from the Deza--Erd\H{o}s--Frankl theorem by a factor of $1-\frac{1}{3r}$. In this article, we improve their result by showing that as $n$ goes to infinity $\Psi_r(n,L)=\Theta_{r,L}(n^{|L|})$ if and only if $\ell_1,\dots,\ell_s,r$ form an arithmetic progression and fully determining the corresponding exact values of $\Psi_r(n,L)$ for sufficiently large $n$ in this case. Moreover, when $L=[t,r-1]$, for the generalized Tur\'an extension of the Erd\H{o}s--Ko--Rado theorem given by Helliar and Liu, we show a Hilton--Milner-type stability result.

math.CO

Monocular Depth Estimation and Segmentation for Transparent Object with Iterative Semantic and Geometric Fusion

Transparent object perception is indispensable for numerous robotic tasks. However, accurately segmenting and estimating the depth of transparent objects remain challenging due to complex optical properties. Existing methods primarily delve into only one task using extra inputs or specialized sensors, neglecting the valuable interactions among tasks and the subsequent refinement process, leading to suboptimal and blurry predictions. To address these issues, we propose a monocular framework, which is the first to excel in both segmentation and depth estimation of transparent objects, with only a single-image input. Specifically, we devise a novel semantic and geometric fusion module, effectively integrating the multi-scale information between tasks. In addition, drawing inspiration from human perception of objects, we further incorporate an iterative strategy, which progressively refines initial features for clearer results. Experiments on two challenging synthetic and real-world datasets demonstrate that our model surpasses state-of-the-art monocular, stereo, and multi-view methods by a large margin of about 38.8%-46.2% with only a single RGB input. Codes and models are publicly available at https://github.com/L-J-Yuan/MODEST.

cs.CV

On low-power error-correcting cooling codes with large distances

A low-power error-correcting cooling (LPECC) code was introduced as a coding scheme for communication over a bus by Chee et al. to control the peak temperature, the average power consumption of on-chip buses, and error-correction for the transmitted information, simultaneously. Specifically, an $(n, t, w, e)$-LPECC code is a coding scheme over $n$ wires that avoids state transitions on the $t$ hottest wires and allows at most $w$ state transitions in each transmission, and can correct up to $e$ transmission errors. In this paper, we study the maximum possible size of an $(n, t, w, e)$-LPECC code, denoted by $C(n,t,w,e)$. When $w=e+2$ is large, we establish a general upper bound $C(n,t,w,w-2)\leq \lfloor \binom{n+1}{2}/\binom{w+t}{2}\rfloor$; when $w=e+2=3$, we prove $C(n,t,3,1) \leq \lfloor \frac{n(n+1)}{6(t+1)}\rfloor$. Both bounds are tight for large $n$ satisfying some divisibility conditions. Previously, tight bounds were known only for $w=e+2=3,4$ and $t\leq 2$. In general, when $w=e+d$ is large for a constant $d$, we determine the asymptotic value of $C(n,t,w,w-d)\sim \binom{n}{d}/\binom{w+t}{d}$ as $n$ goes to infinity, which can be extended to $q$-ary codes.

cs.IT

Mobile Recording Device Recognition Based Cross-Scale and Multi-Level Representation Learning

This paper introduces a modeling approach that employs multi-level global processing, encompassing both short-term frame-level and long-term sample-level feature scales. In the initial stage of shallow feature extraction, various scales are employed to extract multi-level features, including Mel-Frequency Cepstral Coefficients (MFCC) and pre-Fbank log energy spectrum. The construction of the identification network model involves considering the input two-dimensional temporal features from both frame and sample levels. Specifically, the model initially employs one-dimensional convolution-based Convolutional Long Short-Term Memory (ConvLSTM) to fuse spatiotemporal information and extract short-term frame-level features. Subsequently, bidirectional long Short-Term Memory (BiLSTM) is utilized to learn long-term sample-level sequential representations. The transformer encoder then performs cross-scale, multi-level processing on global frame-level and sample-level features, facilitating deep feature representation and fusion at both levels. Finally, recognition results are obtained through Softmax. Our method achieves an impressive 99.6% recognition accuracy on the CCNU_Mobile dataset, exhibiting a notable improvement of 2% to 12% compared to the baseline system. Additionally, we thoroughly investigate the transferability of our model, achieving an 87.9% accuracy in a classification task on a new dataset.

cs.SD

Focal-free uniform hypergraphs and codes

Motivated by the study of a variant of sunflowers, Alon and Holzman recently introduced focal-free hypergraphs. In this paper, we show that there is an interesting connection between the maximum size of focal-free hypergraphs and the renowned Erd\H{o}s Matching Conjecture on the maximum number of edges that can be contained in a uniform hypergraph with bounded matching number. As a consequence, we give asymptotically optimal bounds on the maximum sizes of focal-free uniform hypergraphs and codes, thereby significantly improving the previous results of Alon and Holzman. Moreover, by using the existentce results of combinatorial designs and orthogonal arrays, we are able to explicitly determine the exact sizes of maximum focal-free uniform hypergraphs and codes for a wide range of parameters.

math.CO

Aharonov-Bohm interferometer in inverted-band pn junctions

Inverted-band $pn$ junctions in two-dimensional materials offer a promising platform for electron optics in condensed matter, as they allow to manipulate and guide electron beams without the need for spatial confinement. In this work, we propose the realization of an Aharonov-Bohm (AB) interferometer using such $pn$ junctions. We observe AB oscillations in numerically obtained conductance and analytically identify the conditions for their appearance by analyzing the scattering processes at the $pn$ interface. To support experimental implementation, we also consider junctions with a graded interface, where the potential varies smoothly between the $p$- and $n$-doped regions. Our results reveal an abrupt change in the AB-oscillation frequency, which we attribute to a distinct transition in the hybridization across the interface. We verify that our reported AB oscillations are robust to realistic disorder and temperature decoherence channels. Our study paves the way for realizing the Aharonov-Bohm effect in bulk mesoscopic systems without the need for external spatial confinement, offering new possibilities in electron optics.

cond-mat.mes-hall

Emergent cavity junction around metal-on-graphene contacts

Harnessing graphene devices for applications relies on a comprehensive understanding of how to interact with them. Specifically, scattering processes at the interface with metallic contacts can induce reproducible abnormalities in measurements. Here, we report on emergent transport signatures appearing when contacting sub-micrometer high-quality metallic top contacts to graphene. Using electrostatic simulations and first-principle calculations, we reveal their origin: the contact induces an n-doped radial cavity around it, which is cooperatively defined by the metal-induced electrostatic potential and Klein tunneling. This intricate mechanism leads to secondary resistance peaks as a function of graphene doping that decreases with increasing contact size. Interestingly, in the presence of a perpendicular magnetic field, the cavity spawns a distinct set of Landau levels that interferes with the Landau fan emanating from the graphene bulk. Essentially, an emergent 'second bulk' forms around the contact, as a result of the interplay between the magnetic field and the contact-induced electrostatic potential. The interplay between the intrinsic and emergent bulks leads to direct observation of bulk-boundary correspondence in our experiments. Our work unveils the microscopic mechanisms manifesting at metal-graphene interfaces, opening new avenues for understanding and devising graphene-based electronic devices.

cond-mat.mes-hall

Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization

Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard pseudo-labels including bias accumulation, noise sensitivity, and instability, we propose a novel method named Cross Pseudo-Labeling (XPL), wherein two models learn from each other with the cross-refine mechanism to avoid bias accumulation. We equip XPL with two effective components. Firstly, the soft pseudo-labels with sharpening and pseudo-label exponential moving average mechanisms enable models to achieve gradual self-improvement and ensure stable training. Secondly, the curriculum data selection module adaptively selects pseudo-labels with high quality during training to mitigate potential bias. Experimental results demonstrate that XPL significantly outperforms existing methods, achieving state-of-the-art performance while effectively mitigating confirmation bias and ensuring training stability.

cs.CV

Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise localization, especially for small objects, and suffer from blurry boundaries and false positives. Moreover, the naive semi-supervised method is poor in fully leveraging the information of abundant unlabeled data. In this paper, we propose a novel semi-supervised learning framework for AVSL, namely Dual Mean-Teacher (DMT), comprising two teacher-student structures to circumvent the confirmation bias issue. Specifically, two teachers, pre-trained on limited labeled data, are employed to filter out noisy samples via the consensus between their predictions, and then generate high-quality pseudo-labels by intersecting their confidence maps. The sufficient utilization of both labeled and unlabeled data and the proposed unbiased framework enable DMT to outperform current state-of-the-art methods by a large margin, with CIoU of 90.4% and 48.8% on Flickr-SoundNet and VGG-Sound Source, obtaining 8.9%, 9.6% and 4.6%, 6.4% improvements over self- and semi-supervised methods respectively, given only 3% positional-annotations. We also extend our framework to some existing AVSL methods and consistently boost their performance.

cs.CV

Improved upper bounds for wide-sense frameproof codes

Frameproof codes have been extensively studied for many years due to their application in copyright protection and their connection to extremal set theory. In this paper, we investigate upper bounds on the cardinality of wide-sense $t$-frameproof codes. For $t=2$, we apply results from Sperner theory to give a better upper bound, which significantly improves a recent bound by Zhou and Zhou. For $t\geq 3$, we provide a general upper bound by establishing a relation between wide-sense frameproof codes and cover-free families. Finally, when the code length $n$ is at most $\frac{15+\sqrt{33}}{24}(t-1)^2$, we show that a wide-sense $t$-frameproof code has at most $n$ codewords, and the unique optimal code consists of all weight-one codewords. As byproducts, our results improve several best known results on binary $t$-frameproof codes.

math.CO