SearcharxivSearch

arXiv subjects

Jing He

Publications and source records attributed to Jing He.

At least 19 recordsLinked to original sources

Lightweight Generative Image Semantic Communication over Packet Erasure Channels

This paper addresses packet loss in semantic communication caused by network congestion or channel fluctuations. We propose LGSemCom, a lightweight generative packet-level joint source-channel coding (JSCC) framework for efficient and robust image transmission over packet erasure channels. Unlike conventional distortion-oriented recovery methods that yield blurry averages over erased regions, LGSemCom reformulates image recovery under packet erasures as a conditional generative task. By leveraging adversarial learning, the proposed decoder synthesizes plausible details without increasing complexity during inference. A key innovation of our framework is an erasure-aware weighting (EAW) strategy, which prioritizes generation in erased regions while preserving pixel fidelity in correctly received areas. To ensure computational efficiency, LGSemCom employs a fully convolutional codec based on efficient long-range attention blocks (ELABs) that capture global semantic dependencies with low complexity. Extensive experiments show that LGSemCom achieves superior perceptual reconstruction quality compared with existing benchmarks under severe packet loss, while achieving an order of magnitude faster inference. These attributes make LGSemCom highly suitable for latency-sensitive applications on resource-constrained edge devices.

eess.SP

Deconfining Phase Transition under Real Rotation: A Matrix Model Study

We construct a matrix model to study the deconfining phase transition for a pure gluon plasma that is confined in a cylinder of radius ${\cal R}$ and rotating rigidly at a real-valued angular velocity $\Omega$, satisfying $\mathcal{R} \Omega<1$. The deconfining phase transition arises due to the competition between two terms that constitute the matrix model. The perturbative term comes from the one-loop effective potential computed in the presence of a background field, while the non-perturbative term represents a correction to the perturbative contribution which is brought about by taking into account an effective mass of the gauge fields. Our results show that real rotation induces a radial inhomogeneity of the system and the deconfining temperature $T_c$ drops away from the rotation axis which is consistent with the Tolman-Ehrenfest law. As for the $\Omega$-dependence of $T_c$, it relies on our assumptions of the gluon effective mass. For a constant mass, $T_c$ is found to always decrease with increasing $\Omega$. A non-monotonic behavior of $T_c$ shows up when a $\Omega$-dependent mass is considered, leading to a qualitative change in the region of small angular velocity. In addition, by setting $\Omega=0$ to eliminate rotational effects, we also demonstrate that the finite-volume effect reduces the deconfining temperature relative to the infinite-volume limit. Comparisons between our results and those from various lattice simulations and phenomenological models suggest that controversy remains over how the deconfining phase transition is modified by real rotation and further work is required to reach a definite conclusion.

hep-ph

A negative answer to a question on tilting objects and two-term complexes

Let $M$ be a silting object in an idempotent complete algebraic triangulated category $\mathcal T$. Put $B={\rm End}_{\mathcal T}(M)$, and let $\mathbb{P}_M\colon {\rm pr}(M)\to K^{[-1,0]}({\rm proj}B)$ be the presentation functor associated with $M$. It was recently asked whether $\mathbb{P}_M(T)$ must be tilting whenever $T\in{\rm pr}(M)$ is a tilting object. We answer this question in the negative by giving an explicit finite-dimensional example. Namely, for $$ \Lambda=k(1\xrightarrow{\alpha}2\xrightarrow{\beta}3\xrightarrow{\gamma}4)/(\alpha\beta\gamma),$$ we construct a silting object $M\in K^b({\rm proj}\Lambda)$ and a tilting object $T=\Sigma\Lambda\in{\rm pr}(M)$ for which $$ {\rm dim}_k{\rm Hom}_{K^b({\rm proj}B)}\bigl(\mathbb{P}_M(T),\Sigma^{-1}\mathbb{P}_M(T)\bigr)=1.$$ Thus $\mathbb{P}_M(T)$ is a two-term silting complex but not a tilting complex.

math.RT

A right pretriangulated category which is not right triangulated

Chen, Liu, Lu, and Zhang recently constructed a pretriangulated category with invertible suspension in which Verdier's octahedral axiom fails. We introduce a general enlargement construction for right pretriangulated categories and show that it preserves axioms (RTR1)-(RTR3), while failure of (RTR4) is detected by the forgetful functor. Applied to their type $A_5$ example, the construction yields a right pretriangulated category that is not right triangulated. In this example, the suspension is faithful but not essentially surjective.

math.RT

A pre-$(n+2)$-angulated category which is not $(n+2)$-angulated

We construct an explicit pre-$9$-angulated category which is not $9$-angulated, thereby giving a genuinely higher counterexample to the implication from pre-$(n+2)$-angulated to $(n+2)$-angulated. The underlying additive category is the category of finitely generated projective right modules over the preprojective algebra $\Pi(A_5)$ over $\mathbb F_2$. The construction is obtained by taking an odd power of the twisted complete comparison used by Chen-Liu-Lu-Zhang in their pre-triangulated counterexample and by showing that the resulting pre-$9$-angulation fails the higher mapping-cone axiom.

math.RT

Privacy-Preserving Credit Risk Prediction with Alternative Data

Credit risk prediction is a critical problem in the consumer credit industry. Traditionally, financial institutions construct credit risk prediction models using borrowers' demographic, financial, and credit history data, collectively referred to as traditional data. Recent studies have demonstrated that alternative data, such as borrowers' mobile phone communication data, enable lenders to acquire fuller and more accurate profiles of borrowers' creditworthiness, thereby improving credit risk prediction performance. Nevertheless, alternative data are held by external entities independent of financial institutions. Directly sharing alternative data with financial institutions infringe on consumer privacy, yet existing credit risk prediction studies largely overlook this issue. To address this gap, we define a new problem, namely privacy-preserving credit risk prediction with alternative data, which simultaneously considers three practical constraints: the privacy-preserving constraint that protects consumer privacy, the model-confidentiality constraint that learns and stores the model centrally at the financial institution, and the lossless constraint that maintains the performance of the learned model. To solve this problem, we develop PrivacyCredit, a novel privacy-preserving machine learning method. We then theoretically demonstrate the privacy-preserving, model-confidential, and lossless properties of PrivacyCredit. Through extensive experiments using a real-world credit dataset linked with alternative data, we demonstrate the predictive value of securely incorporating alternative data into credit risk prediction and show that PrivacyCredit achieves the same predictive performance as the model learned from the insecure plaintext combination of traditional and alternative data. We further evaluate its model-confidentiality property and computational efficiency.

cs.LG

Extreme Energy Concentration of Band-Limited Superoscillatory Vortices for Efficient Optical Micromanipulation

The Abbe diffraction limit, tied to the fundamental spatial bandwidth constraint imposed by any physical aperture, remains the primary barrier to achieving ultimate far-field optical resolution and precise light-matter interactions. However, current efforts to engineer structured light fields beyond this limit often come at the cost of massive sacrifices in energy efficiency. In this work, we mathematically complete the family of non-zero azimuthal-order Circular Prolate Spheroidal Wave Functions (CPSWFs), introducing them as a complete class of band-limited superoscillatory optical vortices carrying helical phase. Compared with classical Laguerre-Gaussian (LG) beams, we rigorously prove that these eigenmodes achieve the theoretical upper bound for extreme energy concentration under strict band-limited constraints. At the scale of light-matter interactions, this optimal concentration directly amplifies the intensity gradients and angular momentum densities that govern optical forces. This advantage translates directly into a 29.9% reduction in the trapping power threshold and a 2.3-fold increase in the subdiffraction orbital rotation speed of nanoparticles. Looking forward, this fundamental physical framework not only establishes strict mathematical boundaries for structured light fields but also serves as an absolute theoretical benchmark for deep-learning inverse design, and next-generation extreme optical micro-manipulation systems.

physics.optics

Pareto frontier of portfolio investment under volatility uncertainty and short-sale constraints market

In this paper, we investigate a portfolio investment problem under volatility uncertainty and short-sale constraints market via sublinear expectation which is used to model volatility uncertainty. We assume the stocks admit volatility uncertainty. Thus the related portfolio has upper variance (maximum risk) and lower variance (minimum risk). By introducing a risk factor $w$ to conduct coupled modeling of the maximum and minimum risks, a simplified Sublinear Expectation Mean-Uncertainty Variance (SLE-MUV) model is constructed. Theoretically, we show that the Pareto frontier of the SLE-MUV model is a continuous convex curve, and its optimal solution can be expressed as a polynomial analytical expression with respect to the risk factor $w$. Empirically, we systematically test the practical performance of the SLE-MUV model and conduct comparative analysis with the traditional Mean-Variance (MV) model as the benchmark based on three sets of samples -- simulated generated data, data of the US stock market and the A-share market. The empirical results show that the SLE-MUV model can significantly improving the risk-adjusted return of the investment portfolio.

q-fin.MF

EvoTale: Continual Character Customization for Expanding Story Worlds

Character-centric story visualization aims to synthesize coherent image sequences that depict narrative events and interactions while preserving recurring character identities. In expanding story worlds, new user-specified characters must be continually incorporated despite varying customization difficulty and identity conflicts in multi-character scenes, without disrupting previously learned identities. In this paper, we propose EvoTale, a continual character customization framework for expanding stylized story worlds. We first introduce an All-in-One-World Character Integrator, which accumulates character-specific residual components within a unified LoRA branch using sparsely overlapping subspaces spanned by a shared orthonormal basis, while freezing previously learned components to limit cross-character coupling. We then develop a Character Quality Gate that uses rubric-guided MLLM feedback as a bounded controller to adapt the optimization budget based on the assessed customization quality. Finally, we propose Character-Aware Region-Focus Sampling, which combines bounding-box-guided regional denoising with identity-aware global denoising to preserve character identities within their designated regions while maintaining global narrative coherence. Experimental results show that EvoTale achieves a favorable balance across character fidelity, continual identity retention, multi-character generation quality, and story-text alignment compared with representative story visualization and customization methods.

cs.CV

DVD: Deterministic Video Depth Estimation with Generative Priors

Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand massive labeled datasets to resolve semantic ambiguities. To break this impasse, we present DVD, the first framework to deterministically adapt pre-trained video diffusion models into single-pass depth regressors. Specifically, DVD features three core designs: (i) repurposing the diffusion timestep as a structural anchor to balance global stability with high-frequency details; (ii) latent manifold rectification (LMR) to mitigate regression-induced over-smoothing, enforcing differential constraints to restore sharp boundaries and coherent motion; and (iii) global affine coherence, an inherent property bounding inter-window divergence, which enables seamless long-video inference without requiring complex temporal alignment. Extensive experiments demonstrate that DVD achieves state-of-the-art zero-shot performance across benchmarks. Furthermore, DVD successfully unlocks the profound geometric priors implicit in video foundation models using 163x less task-specific data than leading baselines. Notably, we fully release our pipeline, providing the whole training suite for SOTA video depth estimation to benefit the open-source community.

cs.CV

StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors

The rapid growth of stereoscopic displays, including VR headsets and 3D cinemas, has led to increasing demand for high-quality stereo video content. However, producing 3D videos remains costly and complex, while automatic Monocular-to-Stereo conversion is hindered by the limitations of the multi-stage ``Depth-Warp-Inpaint'' (DWI) pipeline. This paradigm suffers from error propagation, depth ambiguity, and format inconsistency between parallel and converged stereo configurations. To address these challenges, we introduce UniStereo, the first large-scale unified dataset for stereo video conversion, covering both stereo formats to enable fair benchmarking and robust model training. Building upon this dataset, we propose StereoPilot, an efficient feed-forward model that directly synthesizes the target view without relying on explicit depth maps or iterative diffusion sampling. Equipped with a learnable domain switcher and a cycle consistency loss, StereoPilot adapts seamlessly to different stereo formats and achieves improved consistency. Extensive experiments demonstrate that StereoPilot significantly outperforms state-of-the-art methods in both visual fidelity and computational efficiency. Project page: https://hit-perfect.github.io/StereoPilot/.

cs.CV

Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model

Recovering pixel-wise geometric properties from a single image is fundamentally ill-posed due to appearance ambiguity and non-injective mappings between 2D observations and 3D structures. While discriminative regression models achieve strong performance through large-scale supervision, their success is bounded by the scale, quality, and diversity of available data, as well as by limited physical reasoning. Recent diffusion models exhibit powerful world priors that encode geometry and semantics learned from massive image-text data, yet directly reusing their stochastic generative formulation is suboptimal for deterministic geometric inference: the former is optimized for diverse and high-fidelity image generation, whereas the latter requires stable and accurate predictions. In this work, we propose Lotus-2, a two-stage deterministic framework for stable, accurate and fine-grained geometric dense prediction, aiming to provide an optimal adaptation protocol to fully exploit the pre-trained generative priors. Specifically, in the first stage, the core predictor employs a single-step deterministic formulation with a clean-data objective and a lightweight local continuity module (LCM) to generate globally coherent structures without grid artifacts. In the second stage, the detail sharpener performs a constrained multi-step rectified-flow refinement within the manifold defined by the core predictor, enhancing fine-grained geometry through noise-free deterministic flow matching. Using only 59K training samples, less than 1% of existing large-scale datasets, Lotus-2 establishes new state-of-the-art results in monocular depth estimation and highly competitive surface normal prediction. These results demonstrate that diffusion models can serve as deterministic world priors, enabling high-quality geometric reasoning beyond traditional discriminative and generative paradigms.

cs.CV

FactGuard: Event-Centric and Commonsense-Guided Fake News Detection

Fake news detection methods based on writing style have achieved remarkable progress. However, as adversaries increasingly imitate the style of authentic news, the effectiveness of such approaches is gradually diminishing. Recent research has explored incorporating large language models (LLMs) to enhance fake news detection. Yet, despite their transformative potential, LLMs remain an untapped goldmine for fake news detection, with their real-world adoption hampered by shallow functionality exploration, ambiguous usability, and prohibitive inference costs. In this paper, we propose a novel fake news detection framework, dubbed FactGuard, that leverages LLMs to extract event-centric content, thereby reducing the impact of writing style on detection performance. Furthermore, our approach introduces a dynamic usability mechanism that identifies contradictions and ambiguous cases in factual reasoning, adaptively incorporating LLM advice to improve decision reliability. To ensure efficiency and practical deployment, we employ knowledge distillation to derive FactGuard-D, enabling the framework to operate effectively in cold-start and resource-constrained scenarios. Comprehensive experiments on two benchmark datasets demonstrate that our approach consistently outperforms existing methods in both robustness and accuracy, effectively addressing the challenges of style sensitivity and LLM usability in fake news detection.

cs.AI

PDA-LSTM: Knowledge-driven page data arrangement based on LSTM for LCM supression in QLC 3D NAND flash memories

Quarter level cell (QLC) 3D NAND flash memory is emerging as the predominant storage solution in the era of artificial intelligence. QLC 3D NAND flash stores 4 bit per cell to expand the storage density, resulting in narrower read margins. Constrained to read margins, QLC always suffers from lateral charge migration (LCM), which caused by non-uniform charge density across adjacent memory cells. To suppress charge density gap between cells, there are some algorithm in form of intra-page data mapping such as WBVM, DVDS. However, we observe inter-page data arrangements also approach the suppression. Thus, we proposed an intelligent model PDA-LSTM to arrange intra-page data for LCM suppression, which is a physics-knowledge-driven neural network model. PDA-LSTM applies a long-short term memory (LSTM) neural network to compute a data arrangement probability matrix from input page data pattern. The arrangement is to minimize the global impacts derived from the LCM among wordlines. Since each page data can be arranged only once, we design a transformation from output matrix of LSTM network to non-repetitive sequence generation probability matrix to assist training process. The arranged data pattern can decrease the bit error rate (BER) during data retention. In addition, PDA-LSTM do not need extra flag bits to record data transport of 3D NAND flash compared with WBVM, DVDS. The experiment results show that the PDA-LSTM reduces the average BER by 80.4% compared with strategy without data arrangement, and by 18.4%, 15.2% compared respectively with WBVM and DVDS with code-length 64.

cs.AR

DSEBench: A Test Collection for Explainable Dataset Search with Examples

Dataset search is a well-established task in the Semantic Web and information retrieval research. Current approaches retrieve datasets either based on keyword queries or by identifying datasets similar to a given target dataset. These paradigms fail when the information need involves both keywords and target datasets. To address this gap, we investigate a generalized task, Dataset Search with Examples (DSE), and extend it to Explainable DSE (ExDSE), which further requires identifying relevant fields of the retrieved datasets. We construct DSEBench, the first test collection that provides high-quality dataset-level and field-level annotations to support the evaluation of DSE and ExDSE, respectively. In addition, we employ a large language model to generate extensive annotations for training purposes. We establish comprehensive baselines on DSEBench by adapting and evaluating a variety of lexical, dense, and LLM-based retrieval, reranking, and explanation methods.

cs.IR

Early-stopping for Transformer model training

This work, based on Random Matrix Theory (RMT), introduces a novel early-stopping strategy for Transformer training dynamics. Utilizing the Power Law (PL) fit to tansformer attention matrices as a probe, we demarcate training into three stages: structural exploration, heavy-tailed structure stabilization, and convergence saturation. Empirically, we observe that the spectral density of the shallow self-attention matrix $V$ consistently evolves into a heavy-tailed distribution. Crucially, we propose two consistent and validation-set-free criteria: a quantitative metric for heavy-tailed dynamics and a novel spectral signature indicative of convergence. The strong alignment between these criteria highlights the utility of RMT for monitoring and diagnosing the progression of Transformer model training.

cs.LG

DA$^{2}$: Depth Anything in Any Direction

Panorama has a full FoV (360$^\circ\times$180$^\circ$), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimation is gaining increasing traction in 3D vision. However, due to the scarcity of panoramic data, previous methods are often restricted to in-domain settings, leading to poor zero-shot generalization. Furthermore, due to the spherical distortions inherent in panoramas, many approaches rely on perspective splitting (e.g., cubemaps), which leads to suboptimal efficiency. To address these challenges, we propose $\textbf{DA}$$^{\textbf{2}}$: $\textbf{D}$epth $\textbf{A}$nything in $\textbf{A}$ny $\textbf{D}$irection, an accurate, zero-shot generalizable, and fully end-to-end panoramic depth estimator. Specifically, for scaling up panoramic data, we introduce a data curation engine for generating high-quality panoramic depth data from perspective, and create $\sim$543K panoramic RGB-depth pairs, bringing the total to $\sim$607K. To further mitigate the spherical distortions, we present SphereViT, which explicitly leverages spherical coordinates to enforce the spherical geometric consistency in panoramic image features, yielding improved performance. A comprehensive benchmark on multiple datasets clearly demonstrates DA$^{2}$'s SoTA performance, with an average 38% improvement on AbsRel over the strongest zero-shot baseline. Surprisingly, DA$^{2}$ even outperforms prior in-domain methods, highlighting its superior zero-shot generalization. Moreover, as an end-to-end solution, DA$^{2}$ exhibits much higher efficiency over fusion-based approaches. Both the code and the curated panoramic data has be released. Project page: https://depth-any-in-any-dir.github.io/.

cs.CV

SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation

Pre-trained diffusion models provide rich latent features across U-Net levels and are emerging as powerful vision backbones. While prior works such as Marigold and Lotus repurpose diffusion priors for dense geometric perception tasks such as depth and surface normal estimation, their potential for cross-domain human pose estimation remains largely unexplored. Through a systematic analysis of latent features from different upsampling levels of the Stable Diffusion U-Net, we identify the levels that deliver the strongest robustness and cross-domain generalization for pose estimation. Building on these findings, we propose \textbf{SDPose}, which (i) extracts U-Net features from the selected upsampling blocks, (ii) fuses them with a lightweight feature aggregation module to form a robust representation, and (iii) jointly optimizes keypoint heatmap supervision with an auxiliary latent reconstruction loss to regularize training and preserve the pre-trained generative prior. To evaluate cross-domain generalization and robustness, we construct COCO-OOD, a COCO-based benchmark with four subsets: three style-transferred splits to assess domain shift, and one corruption split (noise, weather, digital artifacts, and blur) to test robustness. With a shorter fine-tuning schedule, SDPose achieves performance comparable to Sapiens on COCO, surpasses Sapiens-1B on COCO-WholeBody, and establishes new state-of-the-art results on HumanArt and COCO-OOD.

cs.CV