SearcharxivSearch

arXiv subjects

Yi Ai

Publications and source records attributed to Yi Ai.

7 recordsLinked to original sources

BC-IHV: Conditioning the Color Space for Stable Rectified-Flow Low-Light Enhancement

Low-light image enhancement (LLIE) must correct ambiguous exposure without overwriting structure already supported by the input. Generative transport can model exposure ambiguity; however, its flexibility may also alter observable geometry and chromatic content. Moreover, fixed invertible color coordinates are usually treated only as representations, although their inverse mappings reshape the RGB-domain gradients received by the enhancement network. To address these issues, we propose Structure-Anchored Rectified Flow (SA-RF), which maintains correspondence through separate chromaticity/intensity stems, a scale-matched condition pyramid, and HybridAda. HybridAda assigns location-specific retrieval to spatial cross-attention and global exposure modulation to pooled AdaLN. We further introduce BC-IHV, a learnable Box--Cox polar color space whose analytically invertible intensity mapping controls the inverse-gradient dynamic range through a single exponent. This allows the representation to balance dark-range expansion and gradient conditioning instead of adopting a fixed linear or logarithmic law. Experiments on three LOL benchmarks, blind image-quality evaluation, and cross-dataset tests demonstrate consistent reconstruction and perceptual advantages over the sota. Controlled studies further support the effectiveness of both the proposed framework and color representation.

cs.CV

Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL

Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In enterprise schemas, SQL can execute successfully while using the wrong relationship role or aggregation grain. We study an alternative placement of the stochastic boundary. A multi-turn planner grounds phrases and selects from question-specific governed options; graph traversal, role predicates, grain lowering, SQL construction, and deterministic checks are implemented in code. We evaluate this semantic path compilation (SPC) system against direct DDL-to-SQL generation on the ACME insurance benchmark. On a 38-question adjudicated comparison set with three runs per question, SPC was adjudicated correct on every run for 37 questions (97.4%), compared with 21 (55.3%) for the baseline. The paired discordance was 16 questions in favor of SPC and none in favor of the baseline (two-sided exact McNemar p=3.05x10^-5). SPC answered all 38 questions correctly at least once and produced one refusal and no adjudicated wrong-but-executed run across 114 run outcomes; the baseline produced 29 adjudicated wrong runs and seven additional judge-flagged data-only coincidences on the same set. A strict-equivalence sensitivity analysis increased the paired difference. Additional SPC runs with GPT-5.4 and Gemini-3.6-Flash showed similar question-level robustness, although their per-run verdict artifacts were not preserved. Six additional benchmark items are retained in an all-item analysis and documented separately by failure class. The study supports an end-to-end systems result, not a causal claim that compilation alone produced the gain, because SPC receives governed semantic artifacts that the DDL baseline does not.

cs.DB

FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions

While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encoders struggle with dense spatial tasks due to the loss of visual details caused by low-resolution pretraining and the reliance on noisy, coarse web-crawled image-text pairs. To overcome these limitations, we introduce FineViT, a novel vision encoder specifically designed to unlock fine-grained perception. By replacing coarse web data with dense recaptions, we systematically mitigate information loss through a progressive training paradigm.: first, the encoder is trained from scratch at a high native resolution on billions of global recaptioned image-text pairs, establishing a robust, detail rich semantic foundation. Subsequently, we further enhance its local perception through LLM alignment, utilizing our curated FineCap-450M dataset that comprises over $450$ million high quality local captions. Extensive experiments validate the effectiveness of the progressive strategy. FineViT achieves state-of-the-art zero-shot recognition and retrieval performance, especially in long-context retrieval, and consistently outperforms multimodal visual encoders such as SigLIP2 and Qwen-ViT when integrated into MLLMs. We hope FineViT could serve as a powerful new baseline for fine-grained visual perception.

cs.CV

Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction

Reconstructing a three-dimensional hyperspectral cube from a two-dimensional compressed measurement is a severely ill-posed inverse problem. Existing deep unfolding networks (DUNs) retain fidelity to the imaging model, but regression-trained denoisers can suppress spatial detail and smooth spectral structure under strong modulation. This paper proposes \emph{FMU}, a deep unfolding framework that couples a measurement-conditioned flow-matching prior with a sensing-model-guided measurement update. A two-phase scheme first learns a compact clean-HSI latent target and then trains a conditional velocity field to generate this target from Gaussian noise. A mean-velocity regularizer additionally penalizes the first-moment error of the learned field. On the KAIST 10-scene benchmark under the optical-filter setting, FMU obtains 42.13\,dB PSNR and 0.9900 SSIM, outperforming LADE-DUN by 1.16\,dB in PSNR under the same training data, sensing mask, and evaluation protocol. Under the same optical-filter operator, FMU is further evaluated on held-out KAIST and ICVL scenes without fine-tuning. We also report quantitative simulated-CASSI results and qualitative reconstructions of real CASSI measurements.

cs.CV

Secure Enhancement for RIS-Aided UAV with ISAC: Robust Design and Resource Allocation

This paper analyses the security performance of a reconfigurable intelligent surface (RIS)-aided unmanned aerial vehicle (UAV) communication system with integrated sensing and communications (ISAC). We consider a multiple-antenna UAV transmitting ISAC waveforms to simultaneously detect an untrusted target in the surrounding environment and communicate with a ground Internet-of-Things (IoT) device in the presence of an eavesdropper (Eve). Given that the Eve can conceal their channel state information (CSI) in practical scenarios, we assume that the CSI of the eavesdropper channel is imperfect. For this RIS-aided ISAC-UAV system, we aim to maximize the average communication secrecy rate by jointly optimizing UAV trajectory, RIS passive beamforming, transmit beamforming, and receive beamforming. However, this joint optimization problem is non-convex due to multi-variable coupling. As such, we solve the optimization using an efficient and tractable algorithm using a block coordinate descent (BCD) method. Specifically, we develop a successive convex approximation (SCA) algorithm based on semidefinite relaxation (SDR) to optimise the joint optimization as four separate non-convex subproblems. Numerical results show that our proposed algorithm can successfully ensure the accuracy of sensing targets and significantly improve the communication secrecy rate of the IoT communication devices.

eess.SP

Movable Antenna Enabled ISAC Beamforming Design for Low-Altitude Airborne Vehicles

In mobile systems with low-altitude vehicles, integrated sensing and communication (ISAC) is considered an effective approach to increase the transmission rate due to limited spectrum resources. To further improve the ISAC performance, this paper proposes a novel method called integrated sensing and communication-movable antenna (ISAC-MA) to optimize the antenna's position. Our goal is to support low-space vehicles by optimizing radar and communication joint beamforming and antenna position in the presence of clutter. This scheme not only guarantees the required signal-to-noise ratio (SNR) for sensing but also further improves the SNR for communication. A successive convex approximation (SCA)-based block coordinate descent (BCD) algorithm is proposed to maximize communication capacity under the condition of sensing SNR. Numerical results show that, compared with the traditional ISAC system and various benchmark schemes, the proposed ISAC-MA system can achieve higher communication capacity under the same sensing SNR constraints.

eess.SP

Improving Physical-Layer Security in ISAC-UAV System: Beamforming and Trajectory Optimization

This paper investigates a novel unmanned aerial vehicle (UAV) secure communication system with integrated sensing and communications. We consider wireless security enhancement for a multiple-antenna UAV transmitting ISAC waveforms to communicate with multiple ground Internet-of-Thing devices and detect the surrounding environment. Specifically, we aim to maximize the average communication secrecy rate by optimizing the UAV trajectory and beamforming vectors. Given that the UAV trajectory optimization problem is non-convex due to multi-variable coupling develop an efficient algorithm based on the successive convex approximation (SCA) algorithm. Numerical results show that our proposed algorithm can ensure the accuracy of sensing targets and improve the communication secrecy rate.

eess.SP