SearcharxivSearch

arXiv subjects

Zijian Xu

Publications and source records attributed to Zijian Xu.

13 recordsLinked to original sources

When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation

Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap probe on the encoder's representation promises a way out, but whether it forecasts the expensive outcome has never been tested. We test this with CheapCT on report generation and on MeasureVQA, a new VQA dataset we build. MeasureVQA scores the outcome one capability at a time, its answers measured from segmentation masks and Hounsfield units. Report generation scores the whole report at once and reflects mostly disease. The probe forecasts expensive training across every capability. The rank agreement between probe and fine-tuning stays high throughout, from $ρ=0.90$ to $1.00$. Used to choose an encoder, CheapCT picks one nearly as good as the best while fine-tuning a single candidate, at orders of magnitude less compute. We release the code and MeasureVQA at https://github.com/renjie-liang/CheapCT

cs.CV

ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed before a language model can consume it. Token compression is well studied in general vision, but little of it targets 3D CT specifically. A common baseline is grid average, which pools regular grid cells and can blend distinct anatomy, lesion, and air into one token. We present \textbf{ORCA} (ORgan-Centroid Aggregation), a token compressor for 3D CT. It merges adjacent tokens with organ guidance and adds a sinusoidal encoding of each region's centroid to preserve spatial layout. This preserves the anatomical information a downstream model needs. ORCA is training-free and plug-and-play, producing an adjustable token set without any model change or text query. We evaluate it across two datasets (CT-RATE and Merlin) and five encoders. The evaluation spans two task types: attribute prediction over five families (size, density, location, texture, and disease) and text generation (visual question answering and report generation). At matched token budgets, ORCA improves consistently over existing compression methods. It shrinks the visual context $64\times$ and its KV-cache $50\times$, and is $31\times$ faster to process each volume. Code released at https://github.com/renjie-liang/ORCA-3DCT.

cs.CV

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a newly opened session remains underexplored. We formulate personalized ambiguity adaptation as a new task: given a user's previously resolved coding sessions and a new ambiguous request, an assistant should identify the recurring ambiguity pattern, produce the intended executable solution, and minimize clarification. To benchmark this task, we introduce CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects these mechanisms into unambiguous executable tasks using a controlled three-stage generation pipeline. CAPA contains 600 coding sessions across 60 balanced user--ambiguity cells, including 300 held-out evaluation sessions. We evaluate 12 recent LLMs under no-history and same-user-history conditions using executable success, first-turn success, and turns-to-completion. Our analyses examine task difficulty, user identity, and memory-based history use, and we further propose same-user history gating as a lightweight inference-time method. CAPA provides a foundation for developing long-term coding assistants that better align generated code with user intent while reducing repeated clarification.

cs.AI

Monge solutions and uniqueness in multi-marginal optimal transport with hierarchical jumps

We introduce Hierarchical Jump multi-marginal transport (HJMOT), a generalization of multi-marginal optimal transport where mass can "jump" over intermediate spaces via augmented isolated points. Established on Polish spaces, the framework guarantees the existence of Kantorovich solutions and, under sequential differentiability and a twist condition, the existence and uniqueness of Monge solutions. This core theory extends robustly to diverse settings, including smooth Riemannian manifolds, demonstrating its versatility as a unified framework for optimal transport across complex geometries.

math.PR

Homeomorphism of the Revuz correspondence under Dynkin class assumptions

This paper investigates the topological properties of the Revuz correspondence between positive continuous additive functionals (PCAFs) and their associated smooth measures. Within the Dynkin, local Dynkin, and Green-tight Dynkin classes, we establish bidirectional equivalences among measure convergence, potential convergence, and PCAF convergence. In the local Dynkin class, weak convergence on compact sets, strong $\mathcal{E}_1$-convergence of potentials, uniform convergence of potentials, and $L^1$-convergence of PCAFs are mutually equivalent; under the Green-tight condition, this equivalence extends to the whole space.

math.PR

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native interface for large models, thus requiring additional training of feature projection for semantic alignment. To this end, we propose the multi-scale Gaussian-Language Map (GLMap), which introduces three key designs: (1) explicit geometry, (2) multi-scale semantics covering both instance and region concepts, and (3) a dual-modality interface where each semantic unit jointly stores a natural language description and a 3D Gaussian representation. The 3D Gaussians enable compact storage and fast rendering of task-relevant images via Gaussian splatting. To enable efficient incremental construction, we further propose a Gaussian Estimator that analytically derives Gaussian parameters from dense point clouds without gradient-based optimization. Experiments on ObjectNav, InstNav, and SQA tasks show that GLMap effectively enhances target navigation and contextual reasoning, while remaining compatible with large-model-based methods in a zero-shot manner. The code is available at https://github.com/sx-zhang/GLMap.

cs.CV

DT-BEHRT: Disease Trajectory-aware Transformer for Interpretable Patient Representation Learning

The growing adoption of electronic health record (EHR) systems has provided unprecedented opportunities for predictive modeling to guide clinical decision making. Structured EHRs contain longitudinal observations of patients across hospital visits, where each visit is represented by a set of medical codes. While sequence-based, graph-based, and graph-enhanced sequence approaches have been developed to capture rich code interactions over time or within the same visits, they often overlook the inherent heterogeneous roles of medical codes arising from distinct clinical characteristics and contexts. To this end, in this study we propose the Disease Trajectory-aware Transformer for EHR (DT-BEHRT), a graph-enhanced sequential architecture that disentangles disease trajectories by explicitly modeling diagnosis-centric interactions within organ systems and capturing asynchronous progression patterns. To further enhance the representation robustness, we design a tailored pre-training methodology that combines trajectory-level code masking with ontology-informed ancestor prediction, promoting semantic alignment across multiple modeling modules. Extensive experiments on multiple benchmark datasets demonstrate that DT-BEHRT achieves strong predictive performance and provides interpretable patient representations that align with clinicians' disease-centered reasoning. The source code is publicly accessible at https://github.com/GatorAIM/DT-BEHRT.git.

cs.LG

Stabilizing and Tuning Superconductivity in La$_3$Ni$_2$O$_{7-δ}$ Films: Oxygen Recycling Protocol Reveals Hole-Doping Analogue

The recent achievement of superconductivity in La$_3$Ni$_2$O$_{7-δ}$ with transition temperatures exceeding 40 K in thin films under compressive strain and 80 K in bulk crystals under high pressure opens new avenues for research on high-temperature superconductivity. The realization of superconductivity in thin films requires delicate control of growth conditions, which presents significant challenges in the synthesis process. Furthermore, the stability of superconducting La$_3$Ni$_2$O$_{7-δ}$ films is compromised by oxygen loss, which complicates their characterization. We introduce an effective recycling protocol that involves oxygen removal in a precursor phase followed by ozone-assisted annealing, which restores superconducting properties. By tuning the oxygen content, we construct an electronic phase diagram that highlights oxygen addition as a potential analogue to hole doping via La substitution with Sr, providing insights into the doping mechanism and guiding future material optimization.

cond-mat.supr-con

Large grid subsets without many cospherical points

Motivated by intuitions from projective algebraic geometry, we provide a novel construction of subsets of the $d$-dimensional grid $[n]^d$ of size $n - o(n)$ with no $d + 2$ points on a sphere or a hyperplane. For $d = 2$, this improves the previously best known lower bound of $n/4$ toward the Erdős--Purdy problem due to Thiele in 1995. For $d \ge 3$, this improves the recent $Ω\bigl( n^{\frac{3}{d+1}-o(1)} \bigr)$ bound due to Suk and White, confirming their conjectured $Ω\bigl( n^{\frac{d}{d+1}} \bigr)$ bound in a strong sense, and asymptotically resolves the generalized Erdős--Purdy problem posed by Brass, Moser, and Pach.

math.CO

On $k$-modal subsequences

A $k$-modal sequence is a sequence of real numbers that can be partitioned into $k+1$ (possibly empty) monotone sections such that adjacent sections have opposite monotonicities. For every positive integer $k$, we prove that any sequence of $n$ pairwise distinct real numbers contains a $k$-modal subsequence of length at least $\sqrt{(2k+1)(n-\frac14)} - \frac{k}{2}$, which is tight in a strong sense. This confirms an old conjecture of F.R.K.Chung (J.Combin.Theory Ser.A, 29(3):267-279, 1980).

math.CO

Rainbow even cycles

We prove that every family of (not necessarily distinct) even cycles $D_1, \dotsc, D_{\lfloor 1.2(n-1) \rfloor+1}$ on some fixed $n$-vertex set has a rainbow even cycle (that is, a set of edges from distinct $D_i$'s, forming an even cycle). This resolves an open problem of Aharoni, Briggs, Holzman and Jiang. Moreover, the result is best possible for every positive integer $n$.

math.CO

On the Size of Minimal Separators for Treedepth Decomposition

Treedepth decomposition has several practical applications and can be used to speed up many parameterized algorithms. There are several works aiming to design a scalable algorithm to compute exact treedepth decompositions. Those include works based on a set of all minimal separators. In those algorithms, although a number of minimal separators are enumerated, the minimal separators that are used for an optimal solution are empirically very small. Therefore, analyzing the upper bound on the size of minimal separators is an important problem because it has the potential to significantly reduce the computation time. A minimal separator $S$ is called an optimal top separator if $td(G) = |S| + td(G \backslash S)$, where $td(G)$ denotes the treedepth of $G$. Then, we have two theoretical results on the size of optimal top separators. (1) For any $G$, there is an optimal top separator $S$ such that $|S| \le 2tw(G)$, where $tw(G)$ is the treewidth of $G$. (2) For any $c < 2$, there exists a graph $G$ such that any optimal top separator $S$ of $G$ have $|S| > c \cdot tw(G)$, i.e., the first result gives a tight bound on the size of an optimal top separator.

cs.DS

Solving NP-Hard Problems on Graphs with Extended AlphaGo Zero

There have been increasing challenges to solve combinatorial optimization problems by machine learning. Khalil et al. proposed an end-to-end reinforcement learning framework, S2V-DQN, which automatically learns graph embeddings to construct solutions to a wide range of problems. To improve the generalization ability of their Q-learning method, we propose a novel learning strategy based on AlphaGo Zero which is a Go engine that achieved a superhuman level without the domain knowledge of the game. Our framework is redesigned for combinatorial problems, where the final reward might take any real number instead of a binary response, win/lose. In experiments conducted for five kinds of NP-hard problems including {\sc MinimumVertexCover} and {\sc MaxCut}, our method is shown to generalize better to various graphs than S2V-DQN. Furthermore, our method can be combined with recently-developed graph neural network (GNN) models such as the \emph{Graph Isomorphism Network}, resulting in even better performance. This experiment also gives an interesting insight into a suitable choice of GNN models for each task.

cs.LG