SearcharxivSearch

arXiv subjects

Danyang Chen

Publications and source records attributed to Danyang Chen.

8 recordsLinked to original sources

Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale

Hosted retrieval-augmented generation (RAG) and semantic search allow users to query valuable provider-held corpora, raising two competing demands: to hide each query and chosen result, yet reveal only the documents that the user is authorized to receive. Existing cryptographic approaches either make this costly by processing the entire corpus for every query, or sacrifice quality for efficiency by scanning a few clusters. We repurpose learned deep hashing as a private filter: a randomized binary code points the provider to a short candidate list, while encrypted reranking and oblivious key transfer protect the precise query and final selection. This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents. On the full 2.68M-passage NQ corpus over a 10-Gbps link, our protocol only adds 0.73 seconds, or 10 percent, to a 128-token Qwen3-32B RAG pipeline. The released code satisfies directional metric differential privacy (DP) and substantially reduces embedding-inversion and property-inference leakage, demonstrating that a carefully learned shortlist can make private dense retrieval both accurate and practical.

cs.CR

CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?

Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectures, networking hardware, and distributed communication patterns, making them particularly challenging for code generation models. We present CommBench, a comprehensive benchmark for GPU communication programming, consisting of over 100 expert-curated tasks spanning point-to-point communication, collective operations, expert-parallel communication, compute--communication fusion, and communication utility functions, with reference implementations either written by GPU communication experts or distilled from production codebases. We further introduce a cheat-resistant evaluation framework that automatically compiles, executes, and validates generated code on multi-GPU systems, and a unified metric that jointly measures functional correctness and communication performance. Evaluating leading frontier and open-source code generation models on both intra-node NVLink and inter-node RDMA platforms reveals that even the strongest model, GPT-5.5, correctly implements and achieves competitive performance on only 30.7\% of the benchmark tasks. Our results expose a substantial gap between current LLMs and expert-written GPU communication code, establishing CommBench as a challenging benchmark for advancing AI-assisted systems programming.

cs.DC

Large-scale EM Benchmark for Multi-Organelle Instance Segmentation in the Wild

Accurate instance-level segmentation of organelles in electron microscopy (EM) is critical for quantitative analysis of subcellular morphology and inter-organelle interactions. However, current benchmarks, based on small, curated datasets, fail to capture the inherent heterogeneity and large spatial context of in-the-wild EM data, imposing fundamental limitations on current patch-based methods. To address these limitations, we developed a large-scale, multi-source benchmark for multi-organelle instance segmentation, comprising over 100,000 2D EM images across variety cell types and five organelle classes that capture real-world variability. Dataset annotations were generated by our designed connectivity-aware Label Propagation Algorithm (3D LPA) with expert refinement. We further benchmarked several state-of-the-art models, including U-Net, SAM variants, and Mask2Former. Our results show several limitations: current models struggle to generalize across heterogeneous EM data and perform poorly on organelles with global, distributed morphologies (e.g., Endoplasmic Reticulum). These findings underscore the fundamental mismatch between local-context models and the challenge of modeling long-range structural continuity in the presence of real-world variability. The benchmark dataset and labeling tool will be publicly released soon.

cs.CV

AutoBinder Agent: An MCP-Based Agent for End-to-End Protein Binder Design

Modern AI technologies for drug discovery are distributed across heterogeneous platforms-including web applications, desktop environments, and code libraries-leading to fragmented workflows, inconsistent interfaces, and high integration overhead. We present an agentic end-to-end drug design framework that leverages a Large Language Model (LLM) in conjunction with the Model Context Protocol (MCP) to dynamically coordinate access to biochemical databases, modular toolchains, and task-specific AI models. The system integrates four state-of-the-art components: MaSIF (MaSIF-site and MaSIF-seed-search) for geometric deep learning-based identification of protein-protein interaction (PPI) sites, Rosetta for grafting protein fragments onto protein backbones to form mini proteins, ProteinMPNN for amino acid sequences redesign, and AlphaFold3 for near-experimental accuracy in complex structure prediction. Starting from a target structure, the framework supports de novo binder generation via surface analysis, scaffold grafting and pose construction, sequence optimization, and structure prediction. Additionally, by replacing rigid, script-based workflows with a protocol-driven, LLM-coordinated architecture, the framework improves reproducibility, reduces manual overhead, and ensures extensibility, portability, and auditability across the entire drug design process.

q-bio.BM

Argus: Token Aware Distributed LLM Inference Optimization

Large Language Models (LLMs) are rapidly being integrated into real-world applications, yet their autoregressive architectures introduce significant inference time variability, especially when deployed across heterogeneous edge-cloud systems. Existing solutions largely neglect the dynamic, stochastic, and heterogeneous nature of such environments, often ignoring the impact of variable output token lengths and device diversity. In this work, we present Argus, the first token-aware distributed edge-cloud LLM inference framework that conducts efficient task offloading. Argus features a Length-Aware Semantics (LAS) module, which predicts output token lengths for incoming prompts using a fine-tuned language model with token-length-sensitive feature modulation, enabling precise estimation. Building on this, our Lyapunov-guided Offloading Optimization (LOO) module formulates long-term Quality-of-Experience optimization that explicitly considers both LLM prefilling and decoding costs. We introduce a novel Iterative Offloading Algorithm with Damping and Congestion Control (IODCC) to effectively solve the resulting integer nonlinear programming problem under time-varying constraints. Extensive theoretical and empirical evaluations demonstrate that Argus achieves robust performance and superior efficiency in highly dynamic, heterogeneous settings.

cs.DC

Protomon: A Multimode Qubit in the Fluxonium Molecule

Qubits that are intrinsically insensitive to depolarization and dephasing errors promise to significantly reduce the overhead of fault-tolerant quantum computing. At their optimal operating points, the logical states of these qubits exhibit both exponentially suppressed matrix elements and sweet spots in energy dispersion, rendering the qubits immune to depolarization and dephasing, respectively. We introduce a multimode qubit, the protomon, encoded in a fluxonium molecule circuit. Compared to the closely related $0$-$π$ qubit, the protomon offers several advantages in theory: resilience to circuit parameter disorder, minimal dephasing from intrinsic harmonic modes, and no dependence on static offset charge. As a proof of concept, we realize four protomon qubits. By tuning the qubits to various operating points identified with calibrated two-tone spectroscopy, we measure depolarization times ranging from 64 to 73 $μ$s and dephasing times between 0.2 to 0.5 $μ$s for one selected qubit. The discrepancy between the relatively short measured coherence times and theoretical predictions is not fully understood. This calls for future studies investigating the limiting noise factors, informing the direction for improving coherence times of the protomon qubit.

quant-ph

Tunable inductive coupler for high fidelity gates between fluxonium qubits

The fluxonium qubit is a promising candidate for quantum computation due to its long coherence times and large anharmonicity. We present a tunable coupler that realizes strong inductive coupling between two heavy-fluxonium qubits, each with $\sim50$MHz frequencies and $\sim5$ GHz anharmonicities. The coupler enables the qubits to have a large tuning range of $\textit{XX}$ coupling strengths ($-35$ to $75$ MHz). The $\textit{ZZ}$ coupling strength is $<3$kHz across the entire coupler bias range, and $<100$Hz at the coupler off-position. These qualities lead to fast, high-fidelity single- and two-qubit gates. By driving at the difference frequency of the two qubits, we realize a $\sqrt{i\mathrm{SWAP}}$ gate in $258$ns with fidelity $99.72\%$, and by driving at the sum frequency of the two qubits, we achieve a $\sqrt{b\mathrm{SWAP}}$ gate in $102$ns with fidelity $99.91\%$. This latter gate is only 5 qubit Larmor periods in length. We run cross-entropy benchmarking for over $20$ consecutive hours and measure stable gate fidelities, with $\sqrt{b\mathrm{SWAP}}$ drift ($2 σ$) $< 0.02\%$ and $\sqrt{i\mathrm{SWAP}}$ drift $< 0.08\%$.

quant-ph

Effect of Branching on Phase Behaviors of ABC Triblock Copolymers in Nonfrustrated Systems

The phase behavior of linear dendritic triblock copolymer melts(AB2C4) is studied by self-consistent-field theory (SCFT) in order to find the effects of branching on the phase behavior of ABC linear triblock copolymer melts. We focus on a nonfrustrated parameters that \c{hi}NAB=\c{hi}NBC=\c{hi}NAC=40 where A/C interface will not form. Frank-Kasper phases, asymmetric alternative sphere as well as traditional phases observed in diblock copolymer melts are found to be stable in the calculation.

cond-mat.soft