SearcharxivSearch

arXiv subjects

Qi Wang

Publications and source records attributed to Qi Wang.

At least 37 records · Page 2Linked to original sources

Pressure induced magnetic-field-free superconducting diode effect in NbSe2 flake

The superconducting diode effect (SDE) is a fascinating nonreciprocal phenomenon where the critical current is different for opposite current directions. It is widely believed that realizing SDE requires breaking both inversion symmetry (IS) and time-reversal symmetry (TRS), which are usually achieved via heterostructure engineering and applying external magnetic fields. Here, we report a pressure-induced magnetic-field-free SDE in NbSe2 flakes without any heterostructures. We show that pressure alone breaks the IS, as confirmed by the second harmonic generation. Crucially, upon applying an out-of-plane magnetic field (B), the SDE exhibits even-in-B behavior, implying the absence of explicit TRS breaking. This finding challenges the prevailing theoretical paradigm and demonstrates that a magnetic-field-free SDE can emerge without explicitly breaking TRS. Thereby, our work establishes pressure engineering as a powerful tool for inducing nonreciprocal superconductivity and designing versatile, magnetic-field-free superconducting devices.

cond-mat.supr-con

ProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions

Biological wet-lab protocols are written for trained researchers and often leave routine operations, state-dependent conditions, and contextual parameters implicit, making them difficult to translate into robot-executable actions. We present ProtoAct, a structured protocol-grounding framework that converts free-form biological procedures into state-aware, embodiment-ready action sequences. ProtoAct uses ProtoRAG to retrieve manually annotated examples for context-sensitive parsing, employs RefineChecker to detect and revise missing or inconsistent steps, and applies ActSchema to map the refined procedure into constrained JSON function sequences. We further introduce BioP2E, for which we manually annotate 22 cell-culture protocols into 258 monitoring conditions, 910 executable subtasks, and 962 grounded action calls. Evaluation across seven large language models demonstrates that ProtoAct can be effectively instantiated with different backbones. Ablations confirm that retrieval, posterior checking, and schema constraints make complementary contributions. The parsed subtasks further support demonstration collection and VLA model training, enabling successful execution in both simulation and real-robot settings. ProtoAct thus provides a practical interface between biological protocol understanding and embodied robotic execution.

cs.RO

Towards the Harness of Embodied Agents

The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harness in which an agentic loop orchestrates robot capabilities, each wrapped as a callable tool. It inherits the core components of coding agents, modified as the physical world requires. The world, however, withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. To bridge these gaps, Thea introduces Scene Graph as Context, a persistent, symbolic representation of the world, and Evaluation as Exit Codes, which detects when an action should terminate, judges whether it succeeded, and on failure diagnoses the cause. Together they close the loop between the agent and the physical world. Rich behaviors then emerge from the composition of tools, and the closed loop carries long-horizon tasks to completion in real environments.

cs.AI

Emergence of Double-Dome Superconductivity in the Pressurized Dirac Semimetal BaMg2Bi2

Dirac semimetal BaMg2Bi2 is reported to be a unique topological material that manifests surface superconductivity that coexistswith bulk band topology at ambient pressure. Here, we present a comprehensive investigation of high-pressure superconductingproperties in BaMg2Bi2 single crystal. Significantly, a pressure-driven double-dome superconducting behavior was revealed, withthe superconducting transition temperature Tc approaching the maximum values of 6.67 K at 4.5 GPa and 7.22 K at 10.4 GPafor the first and second superconducting domes, respectively. The combination of high-pressure X-ray diffraction, Hall resistivitymeasurements, and theoretical calculations demonstrates that, the first superconducting regime is closely related to the pressure-modulated Lifshitz transition, whereas the second superconducting phase emerges concurrently with a structural transition fromthe ambient-pressure P3m1 phase to a high-pressure Pnma phase.

cond-mat.supr-con

Pressure-induced concurrent amorphization and superconductivity in topological material NbNiTe5

We have systematically studied the structural and electronic properties of a topological material NbNiTe5 under high pressure. The evolution of the normal state resistance shows a non-monotonic trend from 0.7 GPa to 5.1 GPa, in accordance with the second-order transition along the inter-layer direction observed in X-ray diffraction and Raman spectra. At around 10 GPa, the sample starts amorphization, which is concurrent with the emergence of superconductivity. Upon further compression, the structural disorder enhances and the superconducting transition becomes clearer, suggesting that the superconductivity is modulated by the degree of disorder in NbNiTe5 under high pressure. Within 45.7 GPa, the superconducting transition temperature (Tc) slowly rises from 0.6 K at 9.5 GPa to 1.4 K at 45.7 GPa. Our findings extend the family of transition metal chalcogenide superconductors and shed new light on understanding superconductivity in disordered systems.

cond-mat.supr-con

Superconducting ternary compounds Li-X-B (X=Mo, W) within the mild pressure range: First-principles predictions

Among the superconducting hydrides under high pressure, a number of studies concentrate on the ternary compounds to explore unique superconductors, which are capable of reducing the stable pressure and maintain superconductivity. In this work, to verify our proposed strategy of ternary composition lines (TCLs) to explore ternary compounds, we combined the first-principles calculations and crystal structure predictions to study the ternary compounds Li-X-B (X=Mo, W) under high pressure. After calculations along five and four TCLs in Li-W-B and Li-Mo-B, respectively, five Li-W-B compounds and four Li-Mo-B compounds were predicted. The compositions of LiWB4, Li4MoB2 and LiMo2B2 could be thermodynamically stable under high pressure, and Li2WB6 is around 0.02 eV/atom above the convex hull at 0 GPa, which has potential for synthesizing. Both of the predicted Li2WB6 P6/mmm and Li2WB4 R-3m are superconducting and their Tc are around 11 K, which are similar to the Tc of WB2 P6/mmm around 100 GPa. An anomalous increase of Tc was found in Li4MoB2 C2/m upon compression. We carried out full ternary search (FTS) to evaluate the validity of the TCLs strategy in Li-W-B system at 0 GPa. Our results are helpful for understanding the phase diagram of Li-X-B (X=Mo, W) under high pressure and the introducing of Li atoms provide candidate structures to reduce the measured stable pressure from ~100 GPa in WB2 P6/mmm to 0 GPa. Meanwhile, we preliminary validate the strategy of TCLs in structure predictions and we expect to improve this strategy in the future, shedding light on the studies of ternary compounds.

cond-mat.supr-con

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.

cs.SD

Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images

Full-body capture from unconstrained photographs requires global correspondence across arbitrary views, poses, crops, and occlusions. Yet pose, geometry, and foundation features estimated in this setting are too unreliable for dense matching or appearance transfer, while diffusion rectifiers and optimization pipelines expose no common interface for consuming such uncertain correspondence. Our insight is that correspondence need not be locally accurate: its coarse viewpoint and body layout can still organize how a diffusion prior adapts and guides reconstruction. We introduce \emph{Astrolabe}, a host-portable adapter built on frozen viewpoint-guided spherical maps (SPH). A fixed bounded transform converts SPH into a spatial noise shift, which is matched during prior adaptation and reused during downstream denoising or score-distillation guidance in both pipeline categories. When a rectifier exposes a reference router, the same target/reference SPH additionally supplies coarse compatibility scores to select native appearance features; router-free optimization uses only the shared shift path. Astrolabe therefore follows one SPH--shift--adapt--guide process without dense warping or a learned control branch. Across Puzzle-IOI and 4D-Dress, it improves all reported image metrics in both hosts and all paired Puzzle-IOI geometry metrics; image gains extend to rear views, while 4D-Dress geometry remains stable overall.

cs.CV

On bricks of one-point extension algebras

Let $\Lambda=B[E]$ be the one-point extension algebra of $B$ by the extension module $E$. Using the standard triple description of $\Lambda$-modules, we characterize mixed bricks through an injectivity condition and a scalar-stabilizer condition. Under the assumption that $B$ is $E$-visible finite, we interpret mixed bricks as orbits in Grassmannians and derive a necessary and sufficient criterion for the brick-finiteness of $\Lambda$.

math.RT

AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure

Corporate annual reports contain weakly structured evidence about foreign-exchange risk management, derivative use, natural hedging, and explicit non-use. This study develops AWARE-FX, an auditable AI/NLP decision-support system that converts report text into traceable firm-year hedging-disclosure measures. The system combines a professional-source lexicon, negation and accounting-status logic, channel-specific financial encoders, exact evidence gates, conservative aggregation, and an audit ledger. Across 24,909 Hong Kong firm-years from 2008-2025, it retrieves and scores 543,527 snippets. Reliability is evaluated through ablations, a stratified 300-snippet human audit, three-seed FinBERT-ModernBERT comparisons, strict 2023-2025 temporal tests, probability calibration, selective prediction, and fixed-prompt generative-model benchmarks. FinBERT has the higher mean F1 in seven of eight encoder task-split comparisons; its temporal F1 ranges from 0.702 to 0.872. Abstaining on the 20% least-confident temporal observations raises retained-sample F1 by 0.050-0.077. Deterministic Qwen3-8B performs strongly on commodity and negation evidence but poorly on foreign-debt and accounting-context labels, showing that a general-purpose LLM does not uniformly replace domain constraints. The strict FX score is negatively associated with linked baseline and stress-period FX exposure, whereas the generic broad score is not. These associations provide external construct validation, not causal estimates of hedging effectiveness. AWARE-FX contributes a tested decision-support architecture in which retrieval, status logic, classification, uncertainty handling, aggregation, and external validation remain separately auditable.

cs.CL

CLVisc Agent for autonomous relativistic hydrodynamics studies

We enable large language model (LLM) agents to autonomously perform end-to-end hydrodynamic simulations of the quark-gluon plasma evolution and calculation of final hadron spectra in relativistic heavy-ion collisions. We design a meta skill that allows an agent to explore a project's source code, craft a specialized skill, and iteratively refine it. Applying this meta skill to the (3+1)D viscous hydrodynamic code CLVisc, the agent builds a CLVisc skill encoding its operational knowledge and then independently executes full scientific workflows: designing parameter scans, running simulations, comparing ensemble results, and producing publication-ready figures. Crucially, the agent draws on literature-informed heavy-ion physics to select physically meaningful observables and interpret outcomes without explicit instruction. We demonstrate the pipeline in two scenarios: temperature-dependent shear viscosity over entropy density $\eta/s$, and nuclear-structure effects in O+O collisions at $\sqrt{s_{\mathrm{NN}}} = 5.36$~TeV using four \textit{ab initio} descriptions of $^{16}$O. In both, the agent plans, executes, and analyzes autonomously, devising new initial-state observables to explain final observations and extract qualitative knowledge. The meta skill is agnostic to code versions and Monte Carlo generators, promising future multi-agent systems in high-energy nuclear physics.

nucl-th

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories. BioVLN represents each instrument with three regions: its physical body, a surrounding clearance region, and an operation area in front of the usable side. This model is applied consistently to scene generation, target placement, navigation evaluation, and safety analysis, so success depends on reaching a position from which the instrument can be accessed. BioVLN supports procedural scene generation and manually designed environments, producing 47 scenes and 1667 episodes. Standardized navigation and reinforcement-learning interfaces enable trajectory collection and policy training. Experiments show that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success to 83.3--92.5% and reduces unsafe proximity.

cs.RO

The automorphism groups of random linear codes

The matching codewords framework is a key tool in recent algorithms for solving the Linear Code Equivalence (LCE) problem and in security analyses of LCE-based cryptographic schemes such as LESS. These analyses often rely on the assumption that a random $q$-ary linear code has no monomial automorphisms other than scalar multiples of the identity. For binary codes, Lefmann, Phelps, and R\"odl established the corresponding rigidity phenomenon in the relevant logarithmic dimension range. For general $q$, Hou established an averaged result over all dimensions, whereas the recent prescribed-dimension result of Di Giusto and Ravagnani applies only in a restricted regime near $n/2$. For every fixed prime power $q$ and every fixed real number $\varepsilon>0$, we prove that a uniformly random $k$-dimensional code $\mathcal{C}\subseteq\mathbb{F}_q^n$ has a trivial monomial automorphism group with probability tending to $1$ as $n\to\infty$, provided that $m:=\min\{k,n-k\}\geq(2+\varepsilon)\log_q n$. Furthermore, when $m \le 2 \log_q n + C$, where $C$ is a constant independent of $n$, we also show that the probability that the automorphism group of $\mathcal{C}$ is nontrivial is at least $\frac{1}{2} - \varepsilon$ for large enough $n$.

cs.IT

Texture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion Model

Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore texture maps and focus on natural images. An efficient and generalizable texture super-resolution model can revitalize a large corpus of aging yet valuable assets across industries such as film and video games. We present Texture++, a novel framework for texture super-resolution, which enhances the low-resolution textures of assets to produce high-resolution, high-quality results. Specifically, we reformulate the task of super-resolution in UV space into performing it across multiple rendered views and merging the outputs. Firstly, to achieve more complete and continuous textures in the view space, we propose an adaptive view selection strategy to integrate textures dispersed across UV texture patches. Furthermore, we introduce a quadtree-based texture region organization method for combining super-resolved textures from different viewpoints, providing masks to distinguish regions that require improvement. Finally, we design a diffusion-based super-resolution model that enhances the texture resolution for specified masked regions, seamlessly integrating with surrounding regions. Through comprehensive evaluations, we demonstrate that our approach yields textures with substantially improved detail and coherence over existing methods.

cs.CV

High-accuracy ultrasonic positioning of calibration sources in the Jiangmen Underground Neutrino Observatory

Precise source positioning is essential for detector calibration in large liquid scintillator detectors such as JUNO, particularly in regions where purely mechanical control is insufficient. An ultrasonic positioning system has been developed to reconstruct the three-dimensional coordinates of a calibration source without interfering with photon collection or contaminating the liquid scintillator. The method combines a sound-speed modeling based on dedicated laboratory measurements and in-detector temperature profiles, waveform-based arrival-time reconstruction, and an in-situ calibration of the effective receiver geometry using central-axis deployments. With six active receivers, central-axis positioning yields a mean error of 1.23 cm relative to the known deployment reference. For off-axis operation in the Cable Loop System calibration plane, a detector-realistic simulation that includes timing resolution, sound-speed variation, and receiver-coordinate smearing predicts a positioning uncertainty of 2.40 cm. These results demonstrate that ultrasonic positioning can provide centimetre-level source accuracy for large liquid scintillator detectors and can support off-axis calibration in JUNO-like experiments.

physics.ins-det

Metasurface Antenna-Enabled LEO Satellite Constellation Communications: Design and Optimization

Next-generation low Earth orbit (LEO) satellite constellations face critical bottlenecks in spectral efficiency and onboard hardware complexity. To overcome these limitations, this paper introduces a novel architecture enabled by metasurface antennas (MAs) at the LEO satellites. In particular, MAs are metasurface-integrated feed antennas that perform high-precision beamforming directly in the wave domain, thereby effectively mitigating multi-user interference. Based on such an antenna architecture, a weighted sum rate (WSR) maximization problem is formulated by jointly optimizing the scheduling of feed antennas to terrestrial users (TUs) and the passive beamforming of the metasurface for system performance enhancement. To address this mixed-integer nonlinear programming (MINLP) challenge, an alternating optimization (AO)-based joint scheduling and beamforming algorithm is proposed. On the one hand, the proposed algorithm incorporates a polynomial-time minimum-cost maximum-flow (MCMF) method, which is dedicated to the optimal scheduling of feed antennas and TUs. On the other hand, it adopts a weighted minimum mean square error (WMMSE) method integrated with semidefinite relaxation (SDR) technique, which is tailored for metasurface beamforming design. Simulation results confirm the effectiveness of the proposed algorithm for MA-enabled LEO satellite constellation communications.

cs.IT

Robust Design of Integrated Sensing and Communication in LEO Satellite Systems

With the growing demand for satellite sensing and communication, the limited wireless resources are difficult to support multiple satellite systems. Therefore, it is desired to investigate integrated sensing and communication (ISAC) in low Earth orbit (LEO) satellite systems to enable multi-functionality within a single satellite, thereby saving both spectrum and orbital resources. In this paper, a framework for ISAC in LEO satellite systems is established, where a satellite can simultaneously sense multiple targets and serve multiple communication users (CUs) over the same spectrum. Considering the limited onboard energy of satellite, a novel robust beamforming design algorithm is developed with the goal of minimizing total transmit power while satisfying the mean squared error (MSE) requirements for sensing and signal-to-interference-plus-noise ratio (SINR) requirements for communication in presence of channel phase uncertainty which exacerbates the cross-functional interference. According to theoretical analysis, the proposed algorithm for ISAC in LEO satellite systems is effective. Moreover, extensive simulations confirm the superiority of the proposed algorithm over baselines.

cs.IT

Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational complexity, which leads to substantial inference latency and low control frequency. Asynchronous inference can partially mask this latency by parallelizing action execution and subsequent inference, but it introduces two critical issues: perception-execution misalignment and long reaction time. In this paper, we propose Jetson-PI, a method for efficient VLA deployment on onboard devices via Foresight-Aligned Asynchronous Correction. To address misalignment, we train a lightweight future correction module that predicts future environment representation conditioned on committed actions, enabling the action expert to directly predict actions from the future time step. To reduce reaction time, we introduce confidence-based scheduling optimization that adaptively balances VLM and action expert invocations, complemented by system-level accelerations including CUDA graph reuse, GPU-resident intermediate buffering, and flow unrolling. Extensive experiments demonstrate that Jetson-PI achieves 8.66x and 5.41x improvements in control frequency compared with naive PyTorch and vla.cpp on NVIDIA Jetson Orin, while outperforming VLASH by 14.8\% in average success rate on the LIBERO benchmark. The code of our asynchronous algorithm is available on https://github.com/PKU-SEC-Lab/Jetson-PI, and our efficient llama.cpp-based inference engine is available on https://github.com/PKU-SEC-Lab/Jetson-PI-Edge.

cs.RO