Searcharxiv⌕ Search

arXiv subjects

Yongjie Guan

Publications and source records attributed to Yongjie Guan.

7 recordsLinked to original sources

Bandits with Probing: Optimal Regret and the Limits of Winner Feedback

A learner probes at most $k$ of $n$ arms each round, receives the maximum of their rewards in $[0,1]$, and competes with the best fixed arm. When does the probing advantage pay for learning? We determine two minimax laws. Under independent stochastic rewards with winner feedback (the maximum and a winning label), or on arbitrary fixed sequences given a single signed contrast between block maxima, the minimax regret has order $Φ_{n,k}(T)=\min\{\frac{n-k}{n}T,\frac{n-k}{k}\}$, $2\le k<n$. Under winner feedback, both arbitrary joint i.i.d. rewards and fixed sequences have minimax regret of order $R_{n,k}(T)=\frac{n-k}{n}\min\{T,\frac{n+T}{k},\sqrt{\frac{nT}{k}}\}$. Both laws have universal constants and anytime upper bounds. The first reduces regret to a pure coverage cost: same-round contrasts absorb the stability cost, and independence permits exact resampling whose gains fund sample advancement. The second adds a learning cost that becomes comparable to coverage at horizon $n$; beyond $nk$, numerical maxima improve over labels alone. The lower bound allows every adaptive action size.

cs.LG↗

CavityRank: Zero-Extra-Byte Residual Routing for Cuckoo Filters

Near capacity, a cuckoo filter may reject an insertion even though a legal placement still exists: the table remains structurally feasible, but a bounded policy fails to find an augmenting path. Random kick-out keeps each step cheap but leaves no persistent direction; breadth-first search recovers direction by expanding a frontier and maintaining table-scaled state. CavityRank exploits a resource already present in four-slot packed buckets. Lookup observes only the fingerprint multiset, so query-equivalent lane orders can encode two comparison bits without widening the 64-bit bucket or changing the two-bucket query. The bits form a four-level ordinal residual rank. Insertion follows a minimum-rank edge and re-encodes each modified bucket from its outgoing edges after relocation, propagating the rank actually realized by the packed word. An exact capacity-four orientation oracle separates structural infeasibility from bounded-search loss. In a paired 4,096-bucket XOR16 ladder, CR2 closes 86.47% of Random CF's oracle gap and CavityRank leaves 1.39% of that original gap. A canonical-tie CR2-versus-CavityRank ablation isolates the second implicit bit, which closes 89.95% and 90.69% of CR2's residual gap; the corresponding closures at 65,536 buckets are 83.95% and 84.27%. A separate packed implementation study at 64 MiB and 97.75% load records 42.53 logical reads per insertion, versus 62.67 for explicit labels and 355.74 for depth-10 BFS, with zero extra bytes per bucket and no table-scaled workspace. CavityRank therefore occupies a practical design point between unguided eviction and frontier search.

cs.DS↗

Exact Reachability by Positive Vertex-Centroid Moves

We resolve Problem 60 of The Open Problems Project under the natural convention that a selected vertex at the total centroid is immobile. A cyclically labeled planar $n$-vertex configuration, $n\ge3$, can be transformed exactly into a regular $n$-gon by finitely many vertex-centroid moves if and only if it is noncollinear. The affirmative direction is independent of this convention, and all moves may be chosen positive, so the selected vertex never crosses the centroid of the other vertices. More generally, any two configurations in a connected open subset of the affinely spanning configuration space can be joined by positive moves whose entire continuous execution remains in that subset. For simple labeled $n$-gons, allowing straight-angle vertices, this yields exact reachability while preserving simplicity precisely within each orientation class. In $\mathbb{R}^d$, it yields mutual reachability of all full-dimensional labeled $n$-point configurations when $n\ge d+2$, while orientation is the only obstruction when $n=d+1$. Writing configurations as coordinate matrices, each move becomes a rank-one perturbation of the identity. Explicit one-move curves and three-move conjugation words yield a finite-word endpoint map with invertible differential at every affinely spanning configuration. The inverse function theorem gives exact local reachability with all intermediate states confined to a prescribed neighborhood, and connectedness makes it global. An elementary ear-reduction argument establishes the required connectivity of simple-polygon orientation classes. The same algebra determines exactly the group generated by positive moves. The proof is existential and nonquantitative.

cs.CG↗

AEX: Non-Intrusive Multi-Hop Attestation and Provenance for LLM APIs

Hosted large language models are increasingly accessed through remote APIs, but the API boundary still offers little direct evidence that a returned output actually corresponds to the client-visible request. Recent audits of shadow APIs show that unofficial or intermediary endpoints can diverge from claimed behavior, while existing approaches such as fingerprinting, model-equality testing, verifiable inference, and TEE attestation either remain inferential or answer different questions. We propose AEX, a non-intrusive attestation extension for existing JSON-based LLM APIs. AEX preserves request, response, tool-calling, streaming, and error semantics, and instead adds a signed top-level attestation object that binds a client-visible request projection to either a complete response object or a committed streaming output. To support realistic deployments, AEX provides explicit request-binding modes, signed request-transform receipts for trusted intermediaries, and source-output / output-transform receipts for trusted output rewriting. For streaming, it separates checkpoint proofs for verified prefixes of an unmodified source stream from complete-output lineage for outputs that have been rewritten, buffered, aggregated, or re-packaged, preventing transformed outputs from being mistaken for source-stream prefixes. AEX therefore makes a deliberately narrow claim: a trusted issuer attests to a specific request-output relation, or to a specific complete-output lineage, at the API boundary. We present the protocol design, threat model, verification state machine, security and privacy analysis, an OpenAI-compatible chat-completions profile, and a reference TypeScript prototype with local conformance tests and microbenchmarks.

cs.CR↗

Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates

Consistent hashing is fundamental to distributed systems, but ring-based schemes can exhibit high peak-to-average load ratios unless they use many virtual nodes, while multi-probe methods improve balance at the cost of scattered memory accesses. This paper introduces Local Rendezvous Hashing (LRH), which preserves a token ring but restricts Highest Random Weight (HRW) selection to a cache-local window of C distinct neighboring physical nodes. LRH locates a key by one binary search, enumerates exactly C distinct candidates using precomputed next-distinct offsets, and chooses the HRW winner (optionally weighted). Lookup cost is O(log|R| + C). Under fixed-topology liveness changes, fixed-candidate filtering remaps only keys whose original winner is down, yielding zero excess churn. In a benchmark with N=5000, V=256 (|R|=1.28M), K=50M and C=8, LRH reduces Max/Avg load from 1.2785 to 1.0947 and achieves 60.05 Mkeys/s, about 6.8x faster than multi-probe consistent hashing with 8 probes (8.80 Mkeys/s) while approaching its balance (Max/Avg 1.0697). A microbenchmark indicates multi-probe assignment is dominated by repeated ring searches and memory traffic rather than probe-generation arithmetic.

cs.DC↗

DeepMix: Mobility-aware, Lightweight, and Hybrid 3D Object Detection for Headsets

Mobile headsets should be capable of understanding 3D physical environments to offer a truly immersive experience for augmented/mixed reality (AR/MR). However, their small form-factor and limited computation resources make it extremely challenging to execute in real-time 3D vision algorithms, which are known to be more compute-intensive than their 2D counterparts. In this paper, we propose DeepMix, a mobility-aware, lightweight, and hybrid 3D object detection framework for improving the user experience of AR/MR on mobile headsets. Motivated by our analysis and evaluation of state-of-the-art 3D object detection models, DeepMix intelligently combines edge-assisted 2D object detection and novel, on-device 3D bounding box estimations that leverage depth data captured by headsets. This leads to low end-to-end latency and significantly boosts detection accuracy in mobile scenarios. A unique feature of DeepMix is that it fully exploits the mobility of headsets to fine-tune detection results and boost detection accuracy. To the best of our knowledge, DeepMix is the first 3D object detection that achieves 30 FPS (an end-to-end latency much lower than the 100 ms stringent requirement of interactive AR/MR). We implement a prototype of DeepMix on Microsoft HoloLens and evaluate its performance via both extensive controlled experiments and a user study with 30+ participants. DeepMix not only improves detection accuracy by 9.1--37.3% but also reduces end-to-end latency by 2.68--9.15x, compared to the baseline that uses existing 3D object detection models.

cs.CV↗

DistrEdge: Speeding up Convolutional Neural Network Inference on Distributed Edge Devices

As the number of edge devices with computing resources (e.g., embedded GPUs, mobile phones, and laptops) increases, recent studies demonstrate that it can be beneficial to collaboratively run convolutional neural network (CNN) inference on more than one edge device. However, these studies make strong assumptions on the devices' conditions, and their application is far from practical. In this work, we propose a general method, called DistrEdge, to provide CNN inference distribution strategies in environments with multiple IoT edge devices. By addressing heterogeneity in devices, network conditions, and nonlinear characters of CNN computation, DistrEdge is adaptive to a wide range of cases (e.g., with different network conditions, various device types) using deep reinforcement learning technology. We utilize the latest embedded AI computing devices (e.g., NVIDIA Jetson products) to construct cases of heterogeneous devices' types in the experiment. Based on our evaluations, DistrEdge can properly adjust the distribution strategy according to the devices' computing characters and the network conditions. It achieves 1.1 to 3x speedup compared to state-of-the-art methods.

cs.DC↗