SearcharxivSearch

arXiv subjects

Ya Zhao

Publications and source records attributed to Ya Zhao.

8 recordsLinked to original sources

LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation

Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.

cs.AI

Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition

Cross-lingual Speech Emotion Recognition (CLSER) aims to identify emotional states in unseen languages. However, existing methods heavily rely on the semantic synchrony of complete labels and static feature stability, hindering low-resource languages from reaching high-resource performance. To address this, we propose a semi-supervised framework based on Semantic-Emotional Resonance Embedding (SERE), a cross-lingual dynamic feature paradigm that requires neither target language labels nor translation alignment. Specifically, SERE constructs an emotion-semantic structure using a small number of labeled samples. It learns human emotional experiences through an Instantaneous Resonance Field (IRF), enabling unlabeled samples to self-organize into this structure. This achieves semi-supervised semantic guidance and structural discovery. Additionally, we design a Triple-Resonance Interaction Chain (TRIC) loss to enable the model to reinforce the interaction and embedding capabilities between labeled and unlabeled samples during emotional highlights. Extensive experiments across multiple languages demonstrate the effectiveness of our method, requiring only 5-shot labeling in the source language.

cs.SD

Adjoint SU(5) GUT model with Modular $S_4$ Symmetry

We study the textures of SM fermion mass matrices and their mixings in a supersymmetric adjoint SU(5) Grand Unified Theory with modular $S_4$ being the horizontal symmetry. The Yukawa entries of both quarks and leptons are expressed by modular forms with lower weights. Neutrino sector has an adjoint SU(5) representation 24 as matter superfield, which is a triplet of $S_4$. The effective light neutrino masses is generated through Type-III and Type-I seesaw mechanism. The only common complex parameter in both charged fermion and neutrino sectors is modulus $τ$. Down-type quarks and charged leptons have the same joint effective operators with adjoint scalar in them, and their mass discrepancy in the same generation depends on Clebsch-Gordan factor. Especially for the first two generations the respective Clebsch-Gordan factors made the double Yukawa ratio $y_dy_μ/y_ey_s=12$, in excellent agreement with the experimental result. We reproduce proper CKM mixing parameters and all nine Yukawa eigenvalues of quarks and charged leptons. Neutrino masses and MNS parameters are also produced properly with normal ordering is preferred.

hep-ph

A Cascade Sequence-to-Sequence Model for Chinese Mandarin Lip Reading

Lip reading aims at decoding texts from the movement of a speaker's mouth. In recent years, lip reading methods have made great progress for English, at both word-level and sentence-level. Unlike English, however, Chinese Mandarin is a tone-based language and relies on pitches to distinguish lexical or grammatical meaning, which significantly increases the ambiguity for the lip reading task. In this paper, we propose a Cascade Sequence-to-Sequence Model for Chinese Mandarin (CSSMCM) lip reading, which explicitly models tones when predicting sentence. Tones are modeled based on visual information and syntactic structure, and are used to predict sentence along with visual information and syntactic structure. In order to evaluate CSSMCM, a dataset called CMLR (Chinese Mandarin Lip Reading) is collected and released, consisting of over 100,000 natural sentences from China Network Television website. When trained on CMLR dataset, the proposed CSSMCM surpasses the performance of state-of-the-art lip reading frameworks, which confirms the effectiveness of explicit modeling of tones for Chinese Mandarin lip reading.

cs.CV

Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers

Lip reading has witnessed unparalleled development in recent years thanks to deep learning and the availability of large-scale datasets. Despite the encouraging results achieved, the performance of lip reading, unfortunately, remains inferior to the one of its counterpart speech recognition, due to the ambiguous nature of its actuations that makes it challenging to extract discriminant features from the lip movement videos. In this paper, we propose a new method, termed as Lip by Speech (LIBS), of which the goal is to strengthen lip reading by learning from speech recognizers. The rationale behind our approach is that the features extracted from speech recognizers may provide complementary and discriminant clues, which are formidable to be obtained from the subtle movements of the lips, and consequently facilitate the training of lip readers. This is achieved, specifically, by distilling multi-granularity knowledge from speech recognizers to lip readers. To conduct this cross-modal knowledge distillation, we utilize an efficacious alignment scheme to handle the inconsistent lengths of the audios and videos, as well as an innovative filtering strategy to refine the speech recognizer's prediction. The proposed method achieves the new state-of-the-art performance on the CMLR and LRS2 datasets, outperforming the baseline by a margin of 7.66% and 2.75% in character error rate, respectively.

cs.CV

Micro Congestion Control: Every Flow Deserves a Second Chance

Today, considerable Internet traffic is sent from the datacenter and heads for users. The characteristics of connections served by servers in datacenters are usually diverse and varied over time, with continuous upgrades in network infrastructure and user devices. As a result, a specific congestion control algorithm hardly accommodates the heterogeneity and performs well in various scenarios. In this work, we present Micro Congestion Control (MCC) --- a novel framework for Internet congestion control. With MCC, diverse algorithms can be assigned purposely to connections in one server to adapt to heterogeneity, and different algorithms can be chosen in each connection's life cycle to keep pace with the dynamic of network. We design and implement MCC in Linux, and the experiments validate that MCC is capable of smoothly switching among various candidate algorithms on the fly to achieve potential performance gain in the real world. Meanwhile, the overheads introduced by MCC are moderate and acceptable.

cs.NI

Obtaining nonvanishing $θ_{13}$ with constrained neutrino Yukawa matrix and implications for flavor model buildings

Assuming a diagonal Majorana neutrino mass matrix, we investigate the neutrino Yukawa textures which lead to a non-zero reactor mixing angle $θ_{13}$. The neutrino effective coupling matrix $κ^{eff}$ is pre-diagonalized by a constant mixing pattern $V_ν$ with a vanishing $θ^ν_{13}$. The resulting pre-diagonal symmetrical matrix $κ$ is set to be four texture zeros with two types of off-diagonal elements nonzero, which is $κ_{13}$ and $κ_{23}$, respectively. With the expectation of simple textures we thoroughly classify the linear combinations, $α_{i}$, $β_{i}$ and $γ_{i}$ of Yukawa elements $λ_{ij}$ in a same row, according to the values vanishing or not. Each set of the classifications can lead to a Yukawa texture which may have implications for the discrete flavor model buildings. We also present a model based on $A_{4}$ according to one set of the constraints on the three combinations with a specific choice of a coefficient in Yukawa texture.

hep-ph

SUSY SU(5) $\times S_{4}$ GUT Flavor Model for Fermion Masses and Mixings with Adjoint, Large $θ^{PMNS}_{13}$

We propose an $S_{4}$ flavor model based on supersymmetric (SUSY) SU(5) GUT. The first and third generations of \textbf{10} dimensional representations in SU(5) are all assigned to be $1_{1}$ of $S_{4}$. The second generation of \textbf{10} is to be $1_{2}$ of $S_{4}$. Right-handed neutrinos of singlet \textbf{1} and three generations of $\bar{\textbf{5}}$ are all assigned to be $3_{1}$ of $S_{4}$. The VEVs of two sets of flavon fields are allowed a moderate hierarchy, that is $\langleΦ^ν\rangle \sim λ_{c}\langleΦ^{e}\rangle$. Tri-Bimaximal (TBM) mixing can be produced at both leading order (LO) and next to next to leading order (NNLO) in neutrino sector. All the masses of up-type quarks are obtained at LO. We also get the bottom-tau unification $m_τ=m_{b}$ and the popular Georgi-Jarlskog relation $m_μ=3m_{s}$ as well as a new mass relation $m_{e}=\frac{8}{27}m_{d}$ in which the novel Clebsch-Gordan (CG) factor arises from the adjoint field $H_{24}$. The GUT relation leads to a sizable mixing angle $θ^{e}_{12} \sim θ_{c}$ and the correct quark mixing matrix $V_{CKM}$ can also be realised in the model. The resulting CKM-like mixing matrix of charged leptons modifies the vanishing $θ^ν_{13}$ in TBM mixing to a large $θ^{PMNS}_{13}\simeqθ_{c}/\sqrt{2}$, in excellent agreement with experimental results. A Dirac CP violation phase $ϕ_{12}\simeq\pmπ/2$ is required to make the deviation from $θ^ν_{12}$ small. We also present some phenomenological numerical results predicted by the model.

hep-ph