SearcharxivSearch

arXiv subjects

Xuejun Zhang

Publications and source records attributed to Xuejun Zhang.

14 recordsLinked to original sources

Large-Aperture All-Solid-State Cascaded Liquid-Crystal Beam Steering for High-Resolution Wide-Field Imaging

High-resolution wide-field imaging is essential for applications requiring simultaneous global coverage and local detail, yet conventional approaches face a fundamental trade-off: wide-FOV cameras sacrifice spatial sampling density by distributing finite detector pixels over a broad angular range, while telephoto systems resolve fine features at the cost of scene coverage. Beam-steering devices can mitigate this trade-off but are currently limited in achieving simultaneously all-solid-state, large aperture, and high-speed operation. Here, we report an all-solid-state large-aperture cascaded liquid-crystal beam-steering (CaLiBS) imaging system that extends the effective angular range of a high-resolution narrow-FOV camera by electrically steering sub-FOVs. The CaLiBS module comprises cascaded liquid crystal waveplates and liquid crystal Pancharatnam-Berry phase gratings; a theoretical voltage-prediction model with a hierarchical search algorithm enables efficient calibration under oblique incidence and 10 times faster calibration speed compared with conventional methods. The calibrated system addresses sub-FOVs across 30.3{\deg} * 30.3{\deg} at 2{\deg} intervals with diffraction efficiency above 60%. Sequential sub-FOV acquisition reconstructs a 34.7 * 34.7 composite image, an 8.6-fold enhancement in spatial-bandwidth product over a single-shot wide-FOV camera using the same detector. Combined with object tracking methods, sub-FOV switching further enables high-resolution tracking of moving vehicles within the wide-area scene. This cascaded LC architecture offers a scalable pathway toward compact, vibration-free, and high-resolution wide-field observation.

physics.optics

Predicting Camera Pose from Perspective Descriptions for Spatial Reasoning

Multi-image spatial reasoning remains challenging for current multimodal large language models (MLLMs). While single-view perception is inherently 2D, reasoning over multiple views requires building a coherent scene understanding across viewpoints. In particular, we study perspective taking, where a model must build a coherent 3D understanding from multi-view observations and use it to reason from a new, language-specified viewpoint. We introduce CAMCUE, a pose-aware multi-image framework that uses camera pose as an explicit geometric anchor for cross-view fusion and novel-view reasoning. CAMCUE injects per-view pose into visual tokens, grounds natural-language viewpoint descriptions to a target camera pose, and synthesizes a pose-conditioned imagined target view to support answering. To support this setting, we curate CAMCUE-DATA with 27,668 training and 508 test instances pairing multi-view images and poses with diverse target-viewpoint descriptions and perspective-shift questions. We also include human-annotated viewpoint descriptions in the test split to evaluate generalization to human language. CAMCUE improves overall accuracy by 9.06% and predicts target poses from natural-language viewpoint descriptions with over 90% rotation accuracy within 20{\deg} and translation accuracy within a 0.5 error threshold. This direct grounding avoids expensive test-time search-and-match, reducing inference time from 256.6s to 1.45s per example and enabling fast, interactive use in real-world scenarios.

cs.CV

Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation

Referring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative communication (Grice, 1975). However, current evaluations of vision-language models (VLMs) often overlook the pragmatic dimension, reducing REG to a region-based captioning task and neglecting Gricean maxims. In this work, we revisit REG from a pragmatic perspective, introducing a new dataset (RefOI) of 1.5k images annotated with both written and spoken referring expressions. Through a systematic evaluation of state-of-the-art VLMs, we identify three key failures of pragmatic competence: (1) failure to uniquely identify the referent, (2) inclusion of excessive or irrelevant information, and (3) misalignment with human pragmatic preference, such as the underuse of minimal spatial cues. We also show that standard automatic evaluations fail to capture these pragmatic violations, reinforcing superficial cues rather than genuine referential success. Our findings call for a renewed focus on pragmatically informed models and evaluation frameworks that align with real human communication.

cs.CL

Multi-Object Hallucination in Vision-Language Models

Large vision language models (LVLMs) often suffer from object hallucination, producing objects not present in the given images. While current benchmarks for object hallucination primarily concentrate on the presence of a single object class rather than individual entities, this work systematically investigates multi-object hallucination, examining how models misperceive (e.g., invent nonexistent objects or become distracted) when tasked with focusing on multiple objects simultaneously. We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate ambiguity. With comprehensive empirical studies and analysis of potential factors leading to multi-object hallucination, we found that (1). LVLMs suffer more hallucinations when focusing on multiple objects compared to a single object. (2). The tested object class distribution affects hallucination behaviors, indicating that LVLMs may follow shortcuts and spurious correlations. (3). Hallucinatory behaviors are influenced by data-specific factors, salience and frequency, and model intrinsic behaviors. We hope to enable LVLMs to recognize and reason about multiple objects that often occur in realistic visual scenes, provide insights, and quantify our progress towards mitigating the issues.

cs.CV

Multicolor Ramsey numbers on stars versus pat

For given simple graphs $H_1,H_2,\dots,H_c$, the multicolor Ramsey number $R(H_1,H_2,\dots,H_c)$ is defined as the smallest positive integer $n$ such that for an arbitrary edge-decomposition $\{G_i\}^c_{i=1}$ of the complete graph $K_n$, at least one $G_i$ has a subgraph isomorphic to $H_i$. Let $m,n_1,n_2,\dots,n_c$ be positive integers and $Σ=\sum_{i=1}^{c}(n_i-1)$. Some bounds and exact values of $R(K_{1,n_1},\dots,K_{1,n_c},P_m)$ have been obtained in literature. Wang (Graphs Combin., 2020) conjectured that if $Σ\not\equiv 0\pmod{m-1}$ and $Σ+1\ge (m-3)^2$, then $R(K_{1,n_1},\ldots, K_{1,n_c}, P_m)=Σ+m-1.$ In this note, we give a new lower bound and some exact values of $R(K_{1,n_1},\dots,K_{1,n_c},P_m)$ when $m\leqΣ$, $Σ\equiv k\pmod{m-1}$, and $2\leq k \leq m-2$. These results partially confirm Wang's conjecture.

math.CO

The Tianlin Mission: a 6m UV/Opt/IR space telescope to explore the habitable worlds and the universe

[Abridged] It is expected that the ongoing and future space-borne planet survey missions including TESS, PLATO, and Earth 2.0 will detect thousands of small to medium-sized planets via the transit technique, including over a hundred habitable terrestrial rocky planets. To conduct a detailed study of these terrestrial planets, particularly the cool ones with wide orbits, the exoplanet community has proposed various follow-up missions. The currently proposed ESA mission ARIEL is capable of characterization of planets down to warm super-Earths mainly using transmission spectroscopy. The NASA 6m UV/Opt/NIR mission proposed in the Astro2020 Decadal Survey may further tackle down to habitable rocky planets, and is expected to launch around 2045. In the meanwhile, China is funding a concept study of a 6-m class space telescope named Tianlin (A UV/Opt/NIR Large Aperture Space Telescope) that aims to start its operation within the next 10-15 years and last for 5+ years. Tianlin will be primarily aimed to the discovery and characterization of rocky planets in the habitable zones (HZ) around nearby stars and to search for potential biosignatures mainly using the direct imaging method. Transmission and emission spectroscopy at moderate to high resolution will be carried out as well on a population of exoplanets to strengthen the understanding of the formation and evolution of exoplanets. It will also carry out in-depth studies of the cosmic web and early galaxies, and constrain the nature of the dark matter and dark energy. We describe briefly the primary scientific motivations and main technical considerations based on our preliminary simulation results. We find that a monolithic off-axis space telescope with a primary mirror diameter larger than 6m equipped with a high contrast chronograph can identify water in the atmosphere of a habitable-zone Earth-like planet around a Sun-like star.

astro-ph.EP

Ask Question First for Enhancing Lifelong Language Learning

Lifelong language learning aims to stream learning NLP tasks while retaining knowledge of previous tasks. Previous works based on the language model and following data-free constraint approaches have explored formatting all data as "begin token (\textit{B}) + context (\textit{C}) + question (\textit{Q}) + answer (\textit{A})" for different tasks. However, they still suffer from catastrophic forgetting and are exacerbated when the previous task's pseudo data is insufficient for the following reasons: (1) The model has difficulty generating task-corresponding pseudo data, and (2) \textit{A} is prone to error when \textit{A} and \textit{C} are separated by \textit{Q} because the information of the \textit{C} is diminished before generating \textit{A}. Therefore, we propose the Ask Question First and Replay Question (AQF-RQ), including a novel data format "\textit{BQCA}" and a new training task to train pseudo questions of previous tasks. Experimental results demonstrate that AQF-RQ makes it easier for the model to generate more pseudo data that match corresponding tasks, and is more robust to both sufficient and insufficient pseudo-data when the task boundary is both clear and unclear. AQF-RQ can achieve only 0.36\% lower performance than multi-task learning.

cs.CL

Several Integral Estimates and Some Applications

In this paper, the authors first consider the bidirectional estimates of several typical integrals. As some applications of these integral estimates, the authors investigate the pointwise multipliers from the normal weight general function space $F(p,\mu,s)$ to the normal weight Bloch type space $\mathcal{B_{\nu}}(B_{n})$ on the unit ball $B_{n}$ of $\mathbb{C}^{n}$, where $\mu$ and $\nu$ are two normal functions on $[0,1)$. For the special normal function $\displaystyle{\mu(r)=(1-r^{2})^{\alpha}\log^{\beta}\frac{e}{1-r^{2}}}$ ($\alpha>0$, $-\infty<\beta<\infty$), the authors give the necessary and sufficient conditions of pointwise multipliers from $F(p,\mu,s)$ to $\mathcal{B_{\nu}}(B_{n})$ for all cases.

math.FA

Ces\`{a}ro-like operator acting between Bloch type spaces

Let $\mu$ be a finite positive Borel measure on the interval $[0,1)$ and $f(z)=\sum_{n=0}^{\infty}a_{n}z^{n} \in H(\mathbb{D})$. The Ce\`{a}sro-like operator is defined by $$ \mathcal{C}_\mu(f)(z)=\sum^\infty_{n=0}\mu_n\left(\sum^n_{k=0}a_k\right)z^n, \ z\in \mathbb{D}, $$ where, for $n\geq 0$, $\mu_n$ denotes the $n$-th moment of the measure $\mu$, that is, $\mu_n=\int_{[0, 1)} t^{n}d\mu(t)$. In this paper, we characterize the measures $\mu$ for which $\mathcal{C}_\mu$ is bounded (compact) from one Bloch type space, $\mathcal {B}^{\alpha}$, into another one, $\mathcal {B}^{\beta}$.

math.FA

Generalized integral type Hilbert operator acting on weighted Bloch space

Let $μ$ be a finite Borel measure on $[0,1)$. In this paper, we consider the generalized integral type Hilbert operator $$\mathcal{I}_{μ_{α+1}}(f)(z)=\int_{0}^{1}\frac{f(t)}{(1-tz)^{α+1}}dμ(t)\ \ \ (α>-1).$$ The operator $\mathcal{I}_{μ_{1}}$ has been extensively studied recently. The aim of this paper is to study the boundedness(resp. compactness) of $\mathcal{I}_{μ_{α+1}}$ acting from the normal weight Bloch space into another of the same kind. As consequences of our study, we get completely results for the boundedness of $ \mathcal{I}_{μ_{α+1}}$ acting between Bloch type spaces, logarithmic Bloch spaces among others.

math.FA

Generalized Hilbert operator acting on Bergman spaces

Let $\mu$ be a positive Borel measure on $[0,1)$. If $f \in H(\mathbb{D})$ and $\alpha>-1$, the generalized integral type Hilbert operator defined as follows: $$\mathcal{I}_{\mu_{\alpha+1}}(f)(z)=\int^1_{0} \frac{f(t)}{(1-tz)^{\alpha+1}}d\mu(t), \ \ \ z\in \mathbb{D} .$$ The operator $\mathcal{I}_{\mu_{1}}$ has been extensively studied recently. In this paper, we characterize the measures $\mu$ for which $\mathcal{I}_{\mu_{\alpha+1}}$ is a bounded (resp., compact) operator acting between the Bloch space $\mathcal {B}$ and Bergman space $ A^{p}$, or from $A^{p}(0 -1$.

math.FA

RVAE-LAMOL: Residual Variational Autoencoder to Enhance Lifelong Language Learning

Lifelong Language Learning (LLL) aims to train a neural network to learn a stream of NLP tasks while retaining knowledge from previous tasks. However, previous works which followed data-free constraint still suffer from catastrophic forgetting issue, where the model forgets what it just learned from previous tasks. In order to alleviate catastrophic forgetting, we propose the residual variational autoencoder (RVAE) to enhance LAMOL, a recent LLL model, by mapping different tasks into a limited unified semantic space. In this space, previous tasks are easy to be correct to their own distribution by pseudo samples. Furthermore, we propose an identity task to make the model is discriminative to recognize the sample belonging to which task. For training RVAE-LAMOL better, we propose a novel training scheme Alternate Lag Training. In the experiments, we test RVAE-LAMOL on permutations of three datasets from DecaNLP. The experimental results demonstrate that RVAE-LAMOL outperforms naïve LAMOL on all permutations and generates more meaningful pseudo-samples.

cs.CL

Decomposing Complex Questions Makes Multi-Hop QA Easier and More Interpretable

Multi-hop QA requires the machine to answer complex questions through finding multiple clues and reasoning, and provide explanatory evidence to demonstrate the machine reasoning process. We propose Relation Extractor-Reader and Comparator (RERC), a three-stage framework based on complex question decomposition, which is the first work that the RERC model has been proposed and applied in solving the multi-hop QA challenges. The Relation Extractor decomposes the complex question, and then the Reader answers the sub-questions in turn, and finally the Comparator performs numerical comparison and summarizes all to get the final answer, where the entire process itself constitutes a complete reasoning evidence path. In the 2WikiMultiHopQA dataset, our RERC model has achieved the most advanced performance, with a winning joint F1 score of 53.58 on the leaderboard. All indicators of our RERC are close to human performance, with only 1.95 behind the human level in F1 score of support fact. At the same time, the evidence path provided by our RERC framework has excellent readability and faithfulness.

cs.CL

Reminding the Incremental Language Model via Data-Free Self-Distillation

Incremental language learning with pseudo-data can alleviate catastrophic forgetting in neural networks. However, to obtain better performance, former methods have higher demands for pseudo-data of the previous tasks. The performance dramatically decreases when fewer pseudo-data are employed. In addition, the distribution of pseudo-data gradually deviates from the real data with the sequential learning of different tasks. The deviation will be greater with more tasks learned, which results in more serious catastrophic forgetting. To address these issues, we propose reminding incremental language model via data-free self-distillation (DFSD), which includes self-distillation based on the Earth Mover's Distance and hidden data augmentation. By estimating the knowledge distribution in all layers of GPT-2 and transforming it from teacher model to student model, the Self-distillation based on the Earth Mover's Distance can significantly reduce the demand for pseudo-data. Hidden data augmentation can greatly alleviate the catastrophic forgetting caused by deviations via modeling the generation of pseudo-data as a hidden data augmentation process, where each sample is a mixture of all trained task data. The experimental results demonstrate that our DFSD can exceed the previous state-of-the-art methods even if the maximum decrease in pseudo-data is 90%.

cs.CL