SearcharxivSearch

arXiv subjects

Jialin Wang

Publications and source records attributed to Jialin Wang.

At least 19 recordsLinked to original sources

HierSVA: A Data Synthesis Pipeline, Dataset, and Benchmark for LLM-Driven Hierarchical Hardware Formal Verification

We present HierSVA, an integrated suite that combines a pipeline, dataset, and benchmark for LLM-driven hierarchical hardware formal verification. HierSVA-SP pairs an RTL preprocessing toolchain with an LLM-in-the-loop formal verification flow to produce reference SystemVerilog Assertions (SVA) on hierarchical RTL. Applying it to BaseJump STL yields HierSVA-DS, a dataset of 342 modules, with hierarchy metadata and depths 0--9, accompanied by a deep subset of 28 module-bug pairs with natural-language specifications and bug variants. HierSVA-B decomposes assertion quality into six metric axes: syntax correctness, assertion proof success rate, vacuity, specification faithfulness, mutation coverage, and formal core coverage. Applying HierSVA-B to twelve recent LLMs reveals three findings. First, the module-level compile rate is 67.1\%; among generated assertions in evaluable runs, 82.1\% prove non-vacuously, but the corresponding assertion sets detect only 70.2\% of eligible injected faults and cover 36.2\% of the formal core. Second, on 211 evaluable model--module entries in the deep subset, assertion sets flag buggy RTL with 0.87 recall, but 40\% of predicted-buggy outcomes are false positives on correct RTL, limiting precision to 0.60. Third, agentic mode improves S1-style provability and strength metrics, but gains plateau and oscillate. Codes and artifacts are available at \href{https://github.com/HierSVAAnon/HierSVACodeAndArtifacts}{https://github.com/HierSVAAnon/HierSVACodeAndArtifacts}. Dataset is available at \href{https://huggingface.co/datasets/AnonymousHierSVA/HierSVA}{https://huggingface.co/datasets/AnonymousHierSVA/HierSVA}.

cs.AR

ViraHinter: a dual-modal artificial intelligence framework for predicting virus-host interactions

Protein-protein interactions (PPIs) between a virus and its host govern infection, replication, and pathogenesis. While high-throughput mapping has identified thousands of virus-host associations, much of the virus-host interactome remains uncharacterized due to the labor-intensive nature of experimental screens, the inherent difficulty in capturing transient interactions, and the limited sequence homology across divergent viral families. Here, we introduce ViraHinter, a dual-modal deep learning framework for the precise prediction of virus-host interactions and large-scale inference of interaction landscapes. ViraHinter couples a structure-generation branch with a sequence-representation branch, integrating structure-informed pair representations with ESM-derived embeddings to learn generalizable interaction rules across unseen viruses. We benchmark ViraHinter on pathogenic coronaviruses and influenza A viruses and show that it consistently outperforms RoseTTAFold2-PPI, AlphaFold 3 and RoseTTAFold2-Lite in prioritizing high-confidence candidates even under severe class imbalance and across diverse interface regimes. Notably, it successfully identifies novel functionally relevant host factors and recapitulates the structural plasticity of the complex interfaces. By intersecting predictions across multiple influenza subtypes, ViraHinter reveals 33 shared host factors, offering a roadmap for broad-spectrum antiviral discovery. ViraHinter therefore serves as a robust computational approach for studying virus-host interactions, enabling systematic screening of host factors for all known human-infecting viruses, providing new insights into the shared mechanisms of viral pathogenesis, and accelerating the discovery of novel therapeutic targets and the development of broad-spectrum antivirals.

q-bio.BM

FlexiCamAR: Enhancing Everyday Camera Interactions on AR Glasses with a Flexible Additional Viewpoint

The recent emergence and popularity of consumer-grade augmented reality (AR) glasses from major technology companies highlight their potential to become the next daily computing platform. A dominant design trend in this context is the integration of a front-facing camera to deliver a first-person perspective. While this approach is intuitive, there is limited evidence that it is optimal (or sufficient) for supporting users in daily tasks. This paper explores a more effective camera interaction technique for AR glasses, which we term ``FlexiCamAR." This novel method aims to enhance both efficiency and the range of applications for AR glasses by offering flexible and comfortable secondary camera viewpoints. To investigate the applicability and usability of this approach, we developed a ring camera prototype that can be attached to users' fingers. We then conducted a user study with 12 participants, comparing FlexiCamAR against the baseline, a traditional front-facing AR camera setup, across two common tasks: taking photos and scanning QR codes. Our findings show that FlexiCamAR significantly reduces physical load. We also explore potential scenarios where the additional viewpoint afforded by FlexiCamAR proves valuable, such as capturing low-angle perspectives or navigating confined spaces. Participant feedback further suggests strong potential for additional applications, including selfie taking, video conferencing, and object scanning. Overall, FlexiCamAR presents a novel interaction approach that can serve as a powerful supplement or alternative to the first-person perspective, significantly improving the adaptability of AR glasses for everyday use.

cs.HC

Observable Social Life Spaces: Exploring User Interpretations of agent-side life context in human-agent interaction

Many AI agents are organized around instrumental "command-execution" interactions, where users primarily encounter agents through task requests and responses. Recent work on generative agents and agent life worlds has drawn attention to agents that maintain social contexts beyond direct user commands. In this paper, we study how observable social life spaces shape users' subjective experience, relational interpretations, and perceived equality during human-agent interaction. We introduce the \textit{Observable Social Life Spaces} paradigm, where agents inhabit a continuous virtual environment, engage in daily activities, and form social relationships that users can directly observe. Through an exploratory mixed-methods study ($N=24$), we found that the Observable condition yielded higher perceived-equality ratings and more frequent equality-related role descriptions than the Baseline and Unobservable conditions, but participant-level analysis suggests that the quantitative effect should be interpreted cautiously. We discuss perceived equality as a user-perception signal shaped by this design, with attention to boundary conditions including visual richness, novelty, and person-like attribution from visible agent cues.

cs.HC

The Perceptual Cost of Passthrough: How Video See-Through HMDs Degrade Human Visual Perception of Acuity, Contrast, and Color

Video see-through (VST) technology aims to seamlessly blend the virtual and physical worlds by reconstructing reality through cameras. However, while manufacturers promise high perceptual fidelity, it remains unclear how closely recent commercial VST systems preserve basic visual functions across environmental conditions. In this work, we present an end-to-end perceptual benchmark for three popular VST headsets: Apple Vision Pro, Meta Quest 3, and Meta Quest Pro. Using adapted psychophysical measures, we evaluated participants' visual acuity, contrast sensitivity, and color vision under both normal and low-light conditions, with naked-eye vision as the reference. Our results show measurable gaps between VST and naked-eye performance, especially for visual acuity and contrast sensitivity in low-light environments. By mapping these perceptual gaps across devices, visual functions, and lighting levels, this work provides a practical benchmark for current commercial VST capabilities and highlights where experience design or device optimization may need to compensate for perceptual loss.

cs.HC

Resolution deficits drive simulator sickness and compromise reading performance in virtual environments

Extended reality (XR) is evolving into a general-purpose computing platform, yet its adoption for productivity is hindered by visual fatigue and simulator sickness. While these symptoms are often attributed to latency or motion conflicts, the precise impact of textual clarity on physiological comfort remains undefined. Here we show that sub-optimal effective resolution, the clarity that reaches the eye after the full display-optics-rendering pipeline, is a primary driver of simulator sickness during reading tasks in both virtual reality and video see-through environments. By systematically manipulating end-to-end effective resolution on a unified logMAR scale, we measured reading psychophysics and sickness symptoms in a controlled within-subjects study. We find that reading performance and user comfort degrade exponentially as resolution drops below 0 logMAR (normal visual acuity). Notably, our results reveal 0 logMAR as a key physiological tipping point: resolutions better than this threshold yield naked-eye-level performance with minimal sickness, whereas poorer resolutions trigger rapid, non-linear increases in nausea and oculomotor strain. These findings suggest that the cognitive and perceptual effort required to resolve blurry text directly compromises user comfort, establishing human-eye resolution as a critical baseline for the design of future ergonomic XR systems.

cs.HC

Particle Builder A Board Game for the Teaching of the Standard Model of Particle Physics at a Secondary Level

We present Particle Builder, an online board game which teaches students about concepts from the Standard Model of Particle Physics at a high school level. This short activity resulted in a gain of 0.16, indicating that students learned a significant amount of particle physics knowledge. Students found the activity was more engaging and less difficult than a normal classroom lesson.

physics.ed-ph

Symmetry and uniqueness of the positive solution for the critical Hartree equation on the Heisenberg group

We apply the moving plane method in integral forms to classify the positive solutions of the critical Hartree equation on Heisenberg group \begin{equation}\label{0.1} -\Delta_{\mathbb{H}}u=\left(\int_{\mathbb{H}^{n}}\frac{|u(\xi)|^{Q^{\ast}_{\mu}}}{|\zeta^{-1}\xi|^{\mu}}\mathrm{d}\xi\right)|u|^{Q^{\ast}_{\mu}-2}u,~~~\zeta,\xi\in\mathbb{H}^{n}, \end{equation} where $\Delta_{\mathbb{H}}$ denotes the Kohn Laplacian, $u(\xi)$ is a real-valued function, $Q=2n+2$ is the homogeneous dimension of $\mathbb{H}^{n}$, $\mu\in (0,Q)$ is a real parameter and $Q^{\ast}_{\mu}=\frac{2Q-\mu}{Q-2}$ is the upper critical exponent associated with the Hardy-Littlewood-Sobolev inequality on the Heisenberg group. By introducing the $\mathbb{H}$-reflection, we prove that the solutions of (\ref{0.1}) are cylindrical, upto Heisenberg translation and suitable scaling of function \begin{equation*}\label{0.2} u_{0}(\zeta)=u_{0}(z,t)=\left((1+|z|^{2})^{2}+t^{2}\right)^{-\frac{Q-2}{4}},~~~\zeta=(z,t)\in \mathbb{H}^{n}. \end{equation*} Furthermore, we show that these positive solutions are also CR inversion-symmetric with respect to the unit CC sphere. Consequently, we establish the uniqueness of positive solutions to equation (\ref{0.1}).

math.AP

Semantic-driven Wireless Environment Knowledge Representation for Efficiency-Accuracy Balanced Beam Prediction in Vehicular Networks

The rapid evolution of the internet of vehicles demands ultra-reliable low-latency communication in high-mobility environments, where conventional beam prediction methods suffer from high-dimensional inputs, prolonged training times, and limited interpretability. To address these challenges, the propagation environment semantics-aware wireless environment knowledge beam prediction (PES-WEKBP) framework is proposed. PES-WEKBP pioneers a novel electromagnetic (EM)-grounded knowledge distillation method, transforming raw visual data into an ultra-lean, interpretable material and location-related wireless environment knowledge matrix. This matrix explicitly encodes critical propagation environment semantics, which is material EM properties and spatial relationships through a physics-informed parameterization process, distilling the environment and channel interplay into a minimal yet information-dense representation. A lightweight decision network then leverages this highly compressed knowledge for low-complexity beam prediction. To holistically evaluate the performance of PES-WEKBP, we first design the prediction consistency-efficiency index (PCEI), which combines prediction accuracy with a stability-penalized logarithmic training time to ensure a balanced optimization of reliability and computational efficiency. Experiments validate that PES-WEKBP achieves a 99.75% to 99.96% dimension reduction and improves accuracy by 5.52% to 8.19%, which outperforms state-of-the-art methods in PCEI scores across diverse vehicular scenarios.

eess.SP

Personalized federated prototype learning in mixed heterogeneous data scenarios

Federated learning has received significant attention for its ability to simultaneously protect customer privacy and leverage distributed data from multiple devices for model training. However, conventional approaches often focus on isolated heterogeneous scenarios, resulting in skewed feature distributions or label distributions. Meanwhile, data heterogeneity is actually a key factor in improving model performance. To address this issue, we propose a new approach called PFPL in mixed heterogeneous scenarios. The method provides richer domain knowledge and unbiased convergence targets by constructing personalized, unbiased prototypes for each client. Moreover, in the local update phase, we introduce consistent regularization to align local instances with their personalized prototypes, which significantly improves the convergence of the loss function. Experimental results on Digits and Office Caltech datasets validate the effectiveness of our approach and successfully reduce the communication cost.

cs.LG

Quantitative stability of critical points for the nonlocal-Sobolev inequality in Heisenberg group

We investigate the quantitative stability of the nonlocal Sobolev inequality in Heisenberg group \begin{equation*}\label{non-Sobolev} C_{HL}(Q,\mu) \left(\int_{\mathbb{H}^{n}}\int_{\mathbb{H}^{n}}\frac{|u(\xi)|^{Q^{\ast}_{\mu}}|u(\eta)|^{Q^{\ast}_{\mu}}}{|\eta^{-1}\xi|^{\mu}}\mathrm{d}\xi\mathrm{d}\eta\right)^{\frac{1}{Q^{\ast}_{\mu}}}\leq \int_{\mathbb{H}^{n}}|\nabla_{H}u|^{2}d\xi,\qquad\forall u\in S^{1,2}(\mathbb{H}^{n}), \end{equation*} where $Q=2n+2$ is the homogeneous dimension of the Hiesenberg group $\mathbb{H}^{n}$, $\mu\in(0,Q)$ and $Q^{\ast}_{\mu}=\frac{2Q-\mu}{Q-2}$ are two parameters corresponding to the Hardy-Littlewood-Sobolev inequality and Folland-Stein inequality on Heisenberg group, $C_{HL}(Q,\mu)$ is the sharp constant of the nonlocal-Sobolev inequality. Specifically, when $u$ is close to solving the Euler equation \begin{equation*}\label{non-critical-n} -\Delta_{H} u=\left(\int_{\mathbb{H}^{n}}\frac{|u(\eta)|^{Q^{\ast}_{\mu}}}{|\eta^{-1}\xi|^{\mu}}\mathrm{d}\eta\right)|u|^{Q^{\ast}_{\mu}-2}u,\qquad\xi,\eta\in\mathbb{H}^{n}, \end{equation*} the natural distance between $u$ and the the set of optimizers $U_{\lambda,\zeta}$, defined as $\delta(u)=||\nabla_{H}u-\nabla_{H}U_{\lambda,\zeta}||_{L^{2}}$, can be linearly bounded by the functional derivative term \begin{equation*} \Gamma(u)=\left\|\Delta_{H}u+\left(\int_{\mathbb{H}^{n}}\frac{|u(\eta)|^{Q^{\ast}_{\mu}}}{|\eta^{-1}\xi|^{\mu}}\mathrm{d}\eta\right)|u|^{Q^{\ast}_{\mu}-2}u\right\|_{(S^{1,2}(\mathbb{H}^{n}))^{-1}}. \end{equation*} And for the weakly interacting bubble solutions $\mathop{\sum}\limits_{i=1}^{\nu}U_{\lambda_{i},\zeta_{i}}$, the aforementioned quantitative stability result holds when the dimension $Q=4$.

math.AP

Symmetric powers of $S^{(n-1,1)}$ and $D^{(n-1,1)}$

Let $p$ be a prime and $n\geq 2$ be a positive integer. We establish new formulae for the decompositions of the first $p-1$ symmetric powers of the Specht module $S^{(n-1,1)}$ and the irreducible module $D^{(n-1,1)}$ in characteristic $p$ as direct sums of Young permutation modules. As an application of the formulae, we show that these symmetric powers have Specht filtration and find the vertices of their indecomposable summands. Our main tool, constructed in this paper, is a lift of a splitting map of a short exact sequence of certain symmetric powers to a splitting map of a short exact sequence of higher symmetric powers. This is a general construction, which can be applied to a broader family of modules.

math.RT

Particle Builder -- Learn about the Standard Model while playing against an AI

Particle Builder Online is a web-based education game designed for high school physics students. Students can play against an AI opponent or peers to familiarise themselves with the Standard Model of Particle Physics. The game is aimed at a high school level and tailored to the International Baccalaureate and the Australian Curriculum. Students from four schools in Canberra took pre/post-tests and a survey while completing a lesson where they played Particle Builder. Students' understanding of particle physics concepts improved significantly. Students found the game more enjoyable and effective than regular classroom lessons.

physics.ed-ph

Qwen2.5-Omni Technical Report

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. To enable the streaming of multimodal information inputs, both audio and visual encoders utilize a block-wise processing approach. To synchronize the timestamps of video inputs with audio, we organize the audio and video sequentially in an interleaved manner and propose a novel position embedding approach, named TMRoPE(Time-aligned Multimodal RoPE). To concurrently generate text and speech while avoiding interference between the two modalities, we propose \textbf{Thinker-Talker} architecture. In this framework, Thinker functions as a large language model tasked with text generation, while Talker is a dual-track autoregressive model that directly utilizes the hidden representations from the Thinker to produce audio tokens as output. Both the Thinker and Talker models are designed to be trained and inferred in an end-to-end manner. For decoding audio tokens in a streaming manner, we introduce a sliding-window DiT that restricts the receptive field, aiming to reduce the initial package delay. Qwen2.5-Omni is comparable with the similarly sized Qwen2.5-VL and outperforms Qwen2-Audio. Furthermore, Qwen2.5-Omni achieves state-of-the-art performance on multimodal benchmarks like Omni-Bench. Notably, Qwen2.5-Omni's performance in end-to-end speech instruction following is comparable to its capabilities with text inputs, as evidenced by benchmarks such as MMLU and GSM8K. As for speech generation, Qwen2.5-Omni's streaming Talker outperforms most existing streaming and non-streaming alternatives in robustness and naturalness.

cs.CL

Enhancing Transformer with GNN Structural Knowledge via Distillation: A Novel Approach

Integrating the structural inductive biases of Graph Neural Networks (GNNs) with the global contextual modeling capabilities of Transformers represents a pivotal challenge in graph representation learning. While GNNs excel at capturing localized topological patterns through message-passing mechanisms, their inherent limitations in modeling long-range dependencies and parallelizability hinder their deployment in large-scale scenarios. Conversely, Transformers leverage self-attention mechanisms to achieve global receptive fields but struggle to inherit the intrinsic graph structural priors of GNNs. This paper proposes a novel knowledge distillation framework that systematically transfers multiscale structural knowledge from GNN teacher models to Transformer student models, offering a new perspective on addressing the critical challenges in cross-architectural distillation. The framework effectively bridges the architectural gap between GNNs and Transformers through micro-macro distillation losses and multiscale feature alignment. This work establishes a new paradigm for inheriting graph structural biases in Transformer architectures, with broad application prospects.

cs.LG

Qwen2.5-VL Technical Report

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative functionalities. Qwen2.5-VL achieves a major leap forward in understanding and interacting with the world through enhanced visual recognition, precise object localization, robust document parsing, and long-video comprehension. A standout feature of Qwen2.5-VL is its ability to localize objects using bounding boxes or points accurately. It provides robust structured data extraction from invoices, forms, and tables, as well as detailed analysis of charts, diagrams, and layouts. To handle complex inputs, Qwen2.5-VL introduces dynamic resolution processing and absolute time encoding, enabling it to process images of varying sizes and videos of extended durations (up to hours) with second-level event localization. This allows the model to natively perceive spatial scales and temporal dynamics without relying on traditional normalization techniques. By training a native dynamic-resolution Vision Transformer (ViT) from scratch and incorporating Window Attention, we reduce computational overhead while maintaining native resolution. As a result, Qwen2.5-VL excels not only in static image and document understanding but also as an interactive visual agent capable of reasoning, tool usage, and task execution in real-world scenarios such as operating computers and mobile devices. Qwen2.5-VL is available in three sizes, addressing diverse use cases from edge AI to high-performance computing. The flagship Qwen2.5-VL-72B model matches state-of-the-art models like GPT-4o and Claude 3.5 Sonnet, particularly excelling in document and diagram understanding. Additionally, Qwen2.5-VL maintains robust linguistic performance, preserving the core language competencies of the Qwen2.5 LLM.

cs.CV

Empirical Research on Utilizing LLM-based Agents for Automated Bug Fixing via LangGraph

This paper presents a novel framework for automated code generation and debugging, designed to improve accuracy, efficiency, and scalability in software development. The proposed system integrates three core components LangGraph, GLM4 Flash, and ChromaDB within a four step iterative workflow to deliver robust performance and seamless functionality. LangGraph serves as a graph-based library for orchestrating tasks, providing precise control and execution while maintaining a unified state object for dynamic updates and consistency. It supports multi-agent, hierarchical, and sequential processes, making it highly adaptable to complex software engineering workflows. GLM4 Flash, a large language model, leverages its advanced capabilities in natural language understanding, contextual reasoning, and multilingual support to generate accurate code snippets based on user prompts. ChromaDB acts as a vector database for semantic search and contextual memory storage, enabling the identification of patterns and the generation of context-aware bug fixes based on historical data. The system operates through a structured four-step process: (1) Code Generation, which translates natural language descriptions into executable code; (2) Code Execution, which validates the code by identifying runtime errors and inconsistencies; (3) Code Repair, which iteratively refines buggy code using ChromaDB's memory capabilities and LangGraph's state tracking; and (4) Code Update, which ensures the code meets functional and performance requirements through iterative modifications.

cs.SE