SearcharxivSearch

arXiv subjects

Qi Shen

Publications and source records attributed to Qi Shen.

At least 19 recordsLinked to original sources

Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World

Machine intelligence's push into the physical world is stuck on a gap: deployment demands auditable judgments from day one, fault samples are scarce or absent, and the norms defining "what counts as a fault" live in design documents, not in operational data. We argue this gap is structural, and locate where it can be legitimately closed. We divide the worlds machine intelligence faces into four (phenomenal, basic physical, artificial physical, artificial symbolic) along one axis of constraint strength, and give the Promulgation Criterion: extracting a prior framework from a world is legitimate if and only if the world is intentionally constituted (C1) and has left a readable generative archive (C2). On the criterion's two gradient axes, exactly one world is high on both: the artificial physical world (buildings, factories, infrastructure), whose norms precede their instances; the legitimate path is to extract the framework from the archive, not to induce it from data. We then show what shape such a framework must take: four construction goals force four incompatible carriers, hence at least four layers (syntax, concepts, knowledge, instances); on a closed concept layer fault localization is decidable in polynomial time, and every judgment is interrogable, traceable to a promulgated clause. The same criterion fixes the runtime division of labor with LLMs: promulgatable duties go to rule engines, on-site judgments beyond promulgation go to LLMs, and every generation sandwiched by promulgated clauses is auditable. The theory is falsifiable: four bets (P1-P4) with explicit falsification conditions -- among them that the next large-scale AI breakthrough occurs in the artificial physical world. Evidence: formal proofs (Appendix A); two cases (Appendix B: a cooling plant; the Curiosity rover Sol 1536 anomaly); eight reverse-read lineages, from BACnet to RDF/OWL (Appendix C).

cs.AI

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assignment indirect and obscuring how latent updates shape subsequent reasoning. We introduce GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated continuation. Causal self-attention provides every continuation-token log-probability with a differentiable path to every preceding latent state through the remaining Transformer blocks, enabling reward-weighted gradients from the entire continuation to be assigned directly to the latents. Across five instruction-tuned backbones, three reasoning benchmarks, and two answer formats, GradCuit achieves an average accuracy of 64.5%, outperforming chain-of-thought prompting by 6.6 percentage points and the strongest competing method by 2.4 points. GradCuit also demonstrates greater robustness: across seven learning-rate settings, it consistently outperforms LatentSeek while reducing the standard deviation of accuracy from 1.53 to 0.82, and even its random-walk variant remains competitive with LatentSeek. For interpretability, token-level gradient attribution reveals that latent influence concentrates on reasoning-connector tokens, while layer analysis identifies early-to-middle Transformer layers as the most effective optimization space. By directly optimizing internal reasoning from outcome feedback, GradCuit opens a new axis of robust and interpretable test-time scaling, where LLMs adapt how they reason rather than merely regenerate, sample, or rerank outputs.

cs.LG

BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding

Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one-tower architectures, they still suffer from two primary limitations: (1) over-entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias-aware and reasoning-enhanced framework for one-tower visual grounding. BARE introduces a mechanism that preserves modality-specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state-of-the-art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE.

cs.CV

BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts

Strategic dialogue requires agents to execute distinct dialogue acts, for which belief estimation is essential. While prior work often estimates beliefs accurately, it lacks a principled mechanism to use those beliefs during generation. We bridge this gap by first formalizing two core acts Adversarial and Alignment, and by operationalizing them via probabilistic constraints on what an agent may generate. We instantiate this idea in BEDA, a framework that consists of the world set, the belief estimator for belief estimation, and the conditional generator that selects acts and realizes utterances consistent with the inferred beliefs. Across three settings, Conditional Keeper Burglar (CKBG, adversarial), Mutual Friends (MF, cooperative), and CaSiNo (negotiation), BEDA consistently outperforms strong baselines: on CKBG it improves success rate by at least 5.0 points across backbones and by 20.6 points with GPT-4.1-nano; on Mutual Friends it achieves an average improvement of 9.3 points; and on CaSiNo it achieves the optimal deal relative to all baselines. These results indicate that casting belief estimation as constraints provides a simple, general mechanism for reliable strategic dialogue.

cs.CL

BiHDTrans: binary hyperdimensional transformer for efficient multivariate time series classification

The proliferation of Internet-of-Things (IoT) devices has led to an unprecedented volume of multivariate time series (MTS) data, requiring efficient and accurate processing for timely decision-making in resource-constrained edge environments. Hyperdimensional (HD) computing, with its inherent efficiency and parallelizability, has shown promise in classification tasks but struggles to capture complex temporal patterns, while Transformers excel at sequence modeling but incur high computational and memory overhead. We introduce BiHDTrans, an efficient neurosymbolic binary hyperdimensional Transformer that integrates self-attention into the HD computing paradigm, unifying the representational efficiency of HD computing with the temporal modeling power of Transformers. Empirically, BiHDTrans outperforms state-of-the-art (SOTA) HD computing models by at least 14.47% and achieves 6.67% higher accuracy on average than SOTA binary Transformers. With hardware acceleration on FPGA, our pipelined implementation leverages the independent and identically distributed properties of high-dimensional representations, delivering 39.4 times lower inference latency than SOTA binary Transformers. Theoretical analysis shows that binarizing in holographic high-dimensional space incurs significantly less information distortion than directly binarizing neural networks, explaining BiHDTrans's superior accuracy. Furthermore, dimensionality experiments confirm that BiHDTrans remains competitive even with a 64% reduction in hyperspace dimensionality, surpassing SOTA binary Transformers by 1-2% in accuracy with 4.4 times less model size, as well as further reducing the latency by 49.8% compare to the full-dimensional baseline. Together, these contributions bridge the gap between the expressiveness of Transformers and the efficiency of HD computing, enabling accurate, scalable, and low-latency MTS classification.

cs.LG

Recent advances in DNA origami-engineered nanomaterials and applications

DNA nanotechnology is a unique field, where physics, chemistry, biology, mathematics, engineering, and materials science can elegantly converge. Since the original proposal of Nadrian Seeman, significant advances have been achieved in the past four decades. During this glory time, the DNA origami technique developed by Paul Rothemund further pushed the field forward with a vigorous momentum, fostering a plethora of concepts, models, methodologies, and applications that were not thought of before. This review focuses on the recent progress in DNA origami-engineered nanomaterials in the past five years, outlining the exciting achievements as well as the unexplored research avenues. We believe that the spirits and asset that Seeman left for scientists will continue to bring inter-disciplinary innovations and useful applications to this field in the next decade.

physics.bio-ph

Enhanced timing of a 113 km O-TWTFT link with digital maximum likelihood estimation process

Optical two-way time-frequency transfer (O-TWTFT), employing linear optical sampling and based on frequency combs, is a promising approach for future large-scale optical clock synchronization. It offers the dual benefits of high temporal resolution and an extensive unambiguous range. A critical challenge in establishing long-distance free-space optical links is enhancing detection sensitivity. Particularly at ultra-low received power levels, the error caused by time extraction algorithms for linear optical sampling becomes a significant hindrance to system sensitivity, surpassing the constraints imposed by quantum limitations. In this work, we introduce the Complex Least Squares (CLS) method to enhance both the accuracy and sensitivity of time extraction. Unlike most previous methods that relied solely on phase information, our scheme utilizes a maximum likelihood estimation technique incorporating both amplitude and phase data. Our experiments, conducted over a 113 km free-space link with an average link loss of up to 100 dB, achieved a record minimum received power of 0.1 nW, which is over ten times lower than previous benchmarks. The precision also approaches the quantum limitation.

physics.optics

The row left rank of quaternion unit gain graphs in terms of pendant vertices

Let $\widetilde{G}=(G,U(\mathbb{Q}),\varphi)$ be a quaternion unit gain graph (or $U(\mathbb{Q})$-gain graph), where $G$ is the underlying graph of $\widetilde{G}$, $U(\mathbb{Q})=\{q\in \mathbb{Q}: |q|=1\}$ and $\varphi:\overrightarrow{E}\rightarrow U(\mathbb{Q})$ is the gain function such that $\varphi(e_{ij})=\varphi(e_{ji})^{-1}=\overline{\varphi(e_{ji})}$ for any adjacent vertices $v_{i}$ and $v_{j}$. Let $A(\widetilde{G})$ be the adjacency matrix of $\widetilde{G}$ and let $r(\widetilde{G})$ be the row left rank of $\widetilde{G}$. In this paper, we prove some lower bounds on the row left rank of $U(\mathbb{Q})$-gain graphs in terms of pendant vertices. All corresponding extremal graphs are characterized.

math.CO

The row left rank of a quaternion unit gain graph in terms of maximum degree

Let $\Phi=(G,U(\mathbb{Q}),\varphi)$ be a quaternion unit gain graph (or $U(\mathbb{Q})$-gain graph) of order $n$, $A(\Phi)$ be the adjacency matrix of $\Phi$ and $r(\Phi)$ be the row left rank of $\Phi$. Let $\Delta$ be the maximum degree of $\Phi$. In this paper, we prove that $r(\Phi)\geq\frac{n}{\Delta}$. Moreover, if $\Phi$ is connected, we obtain that $r(\Phi)\geq\frac{n-2}{\Delta-1}$. All the corresponding extremal graphs are characterized.

math.CO

Free-Space Twin-Field Quantum Key Distribution

Twin-field quantum key distribution (TF-QKD) elevates the secure key rate from a linear to a square-root dependence on channel loss while preserving measurement-device-independent security. This protocol is uniquely positioned to enable global-scale quantum networks, even under extreme channel loss. While fiber-based TF-QKD implementations have advanced rapidly since its proposal, free-space realizations have remained elusive due to atmospheric turbulence-induced phase distortions. Here, we report the first experimental demonstration of free-space TF-QKD over 14.2 km urban atmospheric channels, surpassing the effective atmospheric thickness -- a critical threshold for satellite compatibility. We achieve a secret key rate exceeding the repeaterless capacity bound, a milestone for practical quantum communication. Our approach eliminates the need for an auxiliary channel to stabilize a closed interferometer, instead leveraging open-channel time and phase control of optical pulses. This work represents a pivotal advance toward satellite-based global quantum networks, combining high-speed key distribution with inherent resistance to real-world channel fluctuations.

quant-ph

RemiHaven: Integrating "In-Town" and "Out-of-Town" Peers to Provide Personalized Reminiscence Support for Older Drifters

With increasing social mobility and an aging society, more older adults in China are migrating to new cities, known as "older drifters." Due to fewer social connections and cultural adaptation challenges, they face negative emotions such as loneliness and depression. While reminiscence-based interventions have been used to improve older adults' psychological well-being, challenges such as the lack of tangible materials and limited social resources constrain the feasibility of traditional reminiscence approaches for older drifters. To address this challenge, we designed RemiHaven, a personalized reminiscence support tool based on a two-phase formative study. It integrates "In-Town" and "Out-of-Town" peer agents to enhance personalization, engagement, and emotional resonance in the reminiscence process, powered by Multimodal Large Language Models (MLLMs). Our evaluations show RemiHaven's strengths in supporting reminiscence while identifying potential challenges. We conclude by offering insights for the future design of reminiscence support tools for older migrants.

cs.HC

The left row rank of quaternion unit gain graphs in terms of girth

Let $\Phi=(G,U(\mathbb{Q}),\varphi)$ be a quaternion unit gain graph (or $U(\mathbb{Q})$-gain graph). The adjacency matrix of $\Phi$ is denoted by $A(\Phi)$ and the left row rank of $\Phi$ is denoted by $r(\Phi)$. If $\Phi$ has at least one cycle, then the length of the shortest cycle in $\Phi$ is the girth of $\Phi$, denoted by $g$. In this paper, we prove that $r(\Phi)\geq g-2$ for $\Phi$. Moreover, we characterize $U(\mathbb{Q})$-gain graphs satisfy $r(\Phi)=g-i$ ($i=0,1,2$) and all quaternion unit gain graphs with rank 2. The results will generalize the corresponding results of simple graphs (Zhou et al. Linear Algebra Appl. (2021), Duan et al. Linear Algebra Appl. (2024) and Duan, Discrete Math. (2024)), signed graphs (Wu et al. Linear Algebra Appl. (2022)), and complex unit gain graphs (Khan, Linear Algebra Appl. (2024)).

math.CO

113 km absolute ranging with nanometer precision

Accurate long-distance ranging is crucial for diverse applications, including satellite formation flying, very-long-baseline interferometry, gravitational-wave observatory, geographical research, etc. The integration of the time-of-flight mesurement with phase interference in dual-comb method enables high-precision ranging with a rapid update rate and an extended ambiguity range. Pioneering experiments have demonstrated unprecedented precision in ranging, achieving 5 nm @ 60 ms for 1.1 m and 200 nm @ 0.5 s for 25 m. However, long-distance ranging remains technically challenging due to high transmission loss and noise. In this letter, we propose a two-way dual-comb ranging (TWDCR) approach that enables successful ranging over a distance of 113 kilometers. We employ air dispersion analysis and synthetic repetition rate technique to extend the ambiguity range of the inherently noisy channel beyond 100 km. The achieved ranging precision is 11.5 $\mu$m @ 1.3 ms, 681 nm @ 1 s, and 82 nm @ 21 s, as confirmed through a comparative analysis of two independent systems. The advanced long-distance ranging technology is expected to have immediate implications for space research initiatives, such as the space telescope array and the satellite gravimetry.

physics.optics

Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer

In light of the widespread application of Automatic Speech Recognition (ASR) systems, their security concerns have received much more attention than ever before, primarily due to the susceptibility of Deep Neural Networks. Previous studies have illustrated that surreptitiously crafting adversarial perturbations enables the manipulation of speech recognition systems, resulting in the production of malicious commands. These attack methods mostly require adding noise perturbations under $\ell_p$ norm constraints, inevitably leaving behind artifacts of manual modifications. Recent research has alleviated this limitation by manipulating style vectors to synthesize adversarial examples based on Text-to-Speech (TTS) synthesis audio. However, style modifications based on optimization objectives significantly reduce the controllability and editability of audio styles. In this paper, we propose an attack on ASR systems based on user-customized style transfer. We first test the effect of Style Transfer Attack (STA) which combines style transfer and adversarial attack in sequential order. And then, as an improvement, we propose an iterative Style Code Attack (SCA) to maintain audio quality. Experimental results show that our method can meet the need for user-customized styles and achieve a success rate of 82% in attacks, while keeping sound naturalness due to our user study.

cs.SD

Dual-comb spectroscopy over 100km open-air path

Satellite-based greenhouse gases (GHG) sensing technologies play a critical role in the study of global carbon emissions and climate change. However, none of the existing satellite-based GHG sensing technologies can achieve the measurement of broad bandwidth, high temporal-spatial resolution, and high sensitivity at the same time. Recently, dual-comb spectroscopy (DCS) has been proposed as a superior candidate technology for GHG sensing because it can measure broadband spectra with high temporal-spatial resolution and high sensitivity. The main barrier to DCS's display on satellites is its short measurement distance in open air achieved thus far. Prior research has not been able to implement DCS over 20 km of open-air path. Here, by developing a bistatic setup using time-frequency dissemination and high-power optical frequency combs, we have implemented DCS over a 113 km turbulent horizontal open-air path. Our experiment successfully measured GHG with 7 nm spectral bandwidth and a 10 kHz frequency and achieved a CO2 sensing precision of <2 ppm in 5 minutes and <0.6 ppm in 36 minutes. Our results represent a significant step towards advancing the implementation of DCS as a satellite-based technology and improving technologies for GHG monitoring

physics.optics

Text2Bundle: Towards Personalized Query-based Bundle Generation

Bundle generation aims to provide a bundle of items for the user, and has been widely studied and applied on online service platforms. Existing bundle generation methods mainly utilized user's preference from historical interactions in common recommendation paradigm, and ignored the potential textual query which is user's current explicit intention. There can be a scenario in which a user proactively queries a bundle with some natural language description, the system should be able to generate a bundle that exactly matches the user's intention through the user's query and preferences. In this work, we define this user-friendly scenario as Query-based Bundle Generation task and propose a novel framework Text2Bundle that leverages both the user's short-term interests from the query and the user's long-term preferences from the historical interactions. Our framework consists of three modules: (1) a query interest extractor that mines the user's fine-grained interests from the query; (2) a unified state encoder that learns the current bundle context state and the user's preferences based on historical interaction and current query; and (3) a bundle generator that generates personalized and complementary bundles using a reinforcement learning with specifically designed rewards. We conduct extensive experiments on three real-world datasets and demonstrate the effectiveness of our framework compared with several state-of-the-art methods.

cs.IR

Towards Multi-Subsession Conversational Recommendation

Conversational recommendation systems (CRS) could acquire dynamic user preferences towards desired items through multi-round interactive dialogue. Previous CRS mainly focuses on the single conversation (subsession) that user quits after a successful recommendation, neglecting the common scenario where user has multiple conversations (multi-subsession) over a short period. Therefore, we propose a novel conversational recommendation scenario named Multi-Subsession Multi-round Conversational Recommendation (MSMCR), where user would still resort to CRS after several subsessions and might preserve vague interests, and system would proactively ask attributes to activate user interests in the current subsession. To fill the gap in this new CRS scenario, we devise a novel framework called Multi-Subsession Conversational Recommender with Activation Attributes (MSCAA). Specifically, we first develop a context-aware recommendation module, comprehensively modeling user interests from historical interactions, previous subsessions, and feedback in the current subsession. Furthermore, an attribute selection policy module is proposed to learn a flexible strategy for asking appropriate attributes to elicit user interests. Finally, we design a conversation policy module to manage the above two modules to decide actions between asking and recommending. Extensive experiments on four datasets verify the effectiveness of our MSCAA framework for the MSMCR setting.

cs.IR

Video and Audio are Images: A Cross-Modal Mixer for Original Data on Video-Audio Retrieval

Cross-modal retrieval has become popular in recent years, particularly with the rise of multimedia. Generally, the information from each modality exhibits distinct representations and semantic information, which makes feature tends to be in separate latent spaces encoded with dual-tower architecture and makes it difficult to establish semantic relationships between modalities, resulting in poor retrieval performance. To address this issue, we propose a novel framework for cross-modal retrieval which consists of a cross-modal mixer, a masked autoencoder for pre-training, and a cross-modal retriever for downstream tasks.In specific, we first adopt cross-modal mixer and mask modeling to fuse the original modality and eliminate redundancy. Then, an encoder-decoder architecture is applied to achieve a fuse-then-separate task in the pre-training phase.We feed masked fused representations into the encoder and reconstruct them with the decoder, ultimately separating the original data of two modalities. In downstream tasks, we use the pre-trained encoder to build the cross-modal retrieval method. Extensive experiments on 2 real-world datasets show that our approach outperforms previous state-of-the-art methods in video-audio matching tasks, improving retrieval accuracy by up to 2 times. Furthermore, we prove our model performance by transferring it to other downstream tasks as a universal model.

cs.IR