SearcharxivSearch

arXiv subjects

Tian Chen

Publications and source records attributed to Tian Chen.

At least 19 recordsLinked to original sources

Accelerating Data Preprocessing for Efficient Vision Model Inference on Jetson Edge Device

Data preprocessing is a crucial part of deep learning workflows on edge devices. However, decoding data saved in JPEG format is very compute-intensive and occupies a major portion of the preprocessing pipeline. Therefore, increasing the decoding speed is vital for improving overall throughput, especially for inputs with large image sizes, which are often subject to preprocessing bottlenecks. On the other hand, edge devices are equipped with specialized hardware units to accelerate media processing and image decoding. For instance, the NVIDIA Jetson platform possesses a dedicated NVJPEG unit. These units can be used to enhance the performance of the preprocessing pipeline. This paper introduces the utilization of such specific hardware acceleration units for offloading decoding tasks. By combining this with a multi-instance approach, it allows for the parallelization of all compute resources including CPU, NVJPEG, GPU, and DLA in Jetson devices. In this work, we compare various potential pipeline designs. On ResNet18, ResNet50, and ResNet152, three models with different sizes, we evaluate the impact of batch sizes and image sizes, as well as the characteristics of GPU/DLA inference. Finally, a fine-tuning experiment for multi-instance design has been conducted. The multi-instance design with a specific hardware decoding unit involved offers up to 30.02% speedup for large image sizes, compared with the most optimized design without it. Based on these findings, we demonstrate the benefits of using the NVJPEG unit in deep learning workflows and provide guidelines for tuning and optimizing edge inference workflows.

cs.PF

LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.

cs.AI

Radio Map Updating from Streaming Spectrum Measurements via Memory-Based Online Gaussian Processes

Radio maps, which estimate spatial radio-frequency characteristics from spectrum measurements, are essential for applications such as spectrum management and network planning. With the continuous arrival of spectrum measurements, conventional batch processing methods for radio map reconstruction become computationally prohibitive, as they require reprocessing all accumulated measurements for each radio map update. To address this, we propose a memory-based online sparse variational Gaussian process (M-OSVGP) method that efficiently updates radio maps from streaming spectrum measurements. Our method employs sparse variational inference and updates the posterior online by minimizing a hybrid objective that integrates newly received measurements and a memory subset of previous ones to mitigate catastrophic forgetting. To further improve posterior approximation as measurements accumulate over spatially diverse regions, we extend M-OSVGP with a grid-assisted online inducing point selection (GOIPS) algorithm. GOIPS dynamically adapts the number and locations of inducing points based on measurement density and spatial correlation, providing a more informative inducing set while maintaining computational efficiency. Extensive simulations demonstrate the effectiveness of our proposed methods in reconstruction accuracy, computational efficiency, and uncertainty quantification, compared to existing batch and online baselines across various scenarios.

eess.SP

Automatic Model-Hardware Co-Adaptation for Heterogeneous AI Accelerators

Large language models now evolve faster than production inference systems can be ported and optimized. New releases change attention, MoE routing, quantization formats, KV-cache layout, and parallel execution patterns, while deployed accelerator fleets remain heterogeneous across hardware generations, framework forks, operator libraries, compiler backends, and communication runtimes. Serving a new model on existing hardware is therefore a model-framework-kernel-hardware co-adaptation problem. We present MetaInfer, an LLM-agent system that formulates inference adaptation as route search over a costed execution-adaptation graph. The graph connects model semantics, framework dispatch, kernel choices, hardware capabilities, runtime evidence, and serving objectives. MetaInfer constructs and updates this graph during execution, restores missing or blocked routes through patches, and reduces route cost through staged validation and end-to-end profiling. Three real episodes -- DeepSeek V4 Flash on NVIDIA A800, GLM 5.3 Flash on NVIDIA A800, and DeepSeek V4 Flash on Hygon K100AI DCU -- demonstrate deployment repair, cross-model knowledge transfer, and portability across heterogeneous accelerator software stacks.

cs.MA

Risk-sensitive linear-quadratic-Gaussian graphon mean-field games

This paper investigates a class of linear-quadratic-Gaussian risk-sensitive graphon mean-field games, involving an asymptotically infinite population of heterogeneous agents distributed across an asymptotically infinite network, where each agent aims to minimize an exponential cost functional reflecting its risk sensitivity. Following the Nash certainty equivalence methodology, an auxiliary risk-sensitive optimal control problem is constructed and further combined with a consistency condition to determine decentralized strategies of the agents. The well-posedness of the resulting graphon mean-field game equation system, consisting of a family of fully coupled forward-backward differential equations, is established by a fixed point approach under a contraction condition, and by the method of continuity under an operator monotonicity condition, respectively. To prove the epsilon-Nash equilibrium property of the obtained decentralized strategies, one faces significant challenge since the usual L^2 error estimates on mean-field approximations are no longer adequate due to unboundedness of the integrand in the exponentiated cost. The proof will be accomplished by establishing certain exponentiated error estimates instead of L^2 error estimates. Finally, a numerical example is provided to illustrate our results.

math.OC

Entanglement-driven responses through multiscale 3D-printed knits

For their resilience and toughness, filamentous entanglements are ubiquitous in both natural and engineered systems across length scales, from polymer-chain- to collagen-networks and from cable-net structures to forest canopies. Textiles are an everyday manifestation of filamentous entanglement: the remarkable resilience and toughness in knitted fabrics arise predominately from the topology of interlooped yarns. Yet most architected materials do not exploit entanglement as a design primitive, and industrial knitting fixes a narrow set of patterns for manufacturability. Additive manufacturing has recently enabled interlocking structures such as chainmail, knot and woven assemblies, hinting at broader possibilities for entangled architectures. The general challenge is to treat knitting itself as a three-dimensional architected material with predictable and tunable mechanics across scales. Here, we show that knitted architectures fabricated additively can be recast as periodic entangled solids whose responses are both fabric-like and programmable. We reproduce the characteristic behavior of conventional planar knits and extend knitting into the third dimension by interlooping along three orthogonal directions, yielding volumetric knits whose stiffness and dissipation are tuned by prescribed pre-strain. We propose a simple scaling that unifies the responses across stitch geometries and constituent materials. Further, we realize the same topology from centimeter to micrometer scales, culminating in the fabrication of what is, to our knowledge, the smallest knitted structure ever made. By demonstrating 3D-printed knits can be interpreted both as a traditional fabric, as well as a novel architected material with defined periodicity, this work establishes the dual nature of entangled filaments and paves the way towards a new form of material architectures with high degrees of entanglement.

physics.app-ph

A Continuum Macro-Model for Bistable Periodic Auxetic Surfaces

A macro-constitutive model for the deformation response of periodic rotating bistable auxetic surfaces is developed. Focus is placed on isotropic surfaces made of bistable hexagonal cells composed of six triangular units with two stable equilibrium states. Adopting a variational formulation, the effective stress-strain response is derived from a free energy function expressed in terms of the invariants of the logarithmic strain. To address the mathematical ill-posedness and numerical artifacts--such as mesh sensitivity--arising from the double-well nature of the free energy, two regularization approaches are introduced: (i) a gradient-enhanced first invariant of the logarithmic strain, and (ii) an artificial material rate dependency. Although neither regularization guarantees solution uniqueness, the former mitigates mesh sensitivity, while the latter improves the convergence behavior of the nonlinear numerical scheme by promoting smooth temporal evolution of transition localization and enabling the system to overcome snap-backs induced by local non-proportional loading near transition fronts. The model is implemented using membrane/shell structural elements and plane stress continuum ones within the ABAQUS finite element suite. Numerical simulations demonstrate the efficacy of the proposed formulation and its implementation.

physics.app-ph

Indefinite Linear-Quadratic Partially Observed Mean-Field Game

This paper investigates an indefinite linear-quadratic partially observed mean-field game with common noise, incorporating both state-average and control-average effects. In our model, each agent's state is observed through both individual and public observations, which are modeled as general stochastic processes rather than Brownian motions. {It is noteworthy that} the weighting matrices in the cost functional are allowed to be indefinite. We derive the optimal decentralized strategies using the Hamiltonian approach and establish the well-posedness of the resulting Hamiltonian system by employing a relaxed compensator. The associated consistency condition and the feedback representation of decentralized strategies are also established. Furthermore, we demonstrate that the set of decentralized strategies form an $\varepsilon$-Nash equilibrium. As an application, we solve a mean-variance portfolio selection problem.

math.OC

Acoustic-Assisted Fabrication of Thin Shells with Spatially Distributed Imperfections

Thin-shell structures, found in biological systems such as beetle carapaces and widely used in aerospace, civil, and mechanical engineering, achieve remarkable strength-to-mass ratio given their slenderness and curved geometries. However, their load-bearing capacity is highly sensitive to geometric imperfections, which are often unavoidable during fabrication and can trigger subcritical buckling. Silicone-based hemispherical domes have served as an experimental modal system to study this phenomenon, yet prior work has largely focused on localized dimples or flat imperfections, failing to capture the spatially distributed nature of real-world imperfection patterns. Here, we introduce an acoustic-assisted method for fabricating thin shells with spatially distributed, vibrational mode-shaped imperfections. Silicone is cast onto a thick elastic mold excited by a speaker, and vibration-induced flow during curing creates thickness variations. High-speed imaging and microCT scanning reveal accumulation of material at the antinodes of the mold's vibrational modes. We show that imperfection geometry can be tuned by acoustic frequency, while their amplitude increases with acoustic volume. Buckling experiments demonstrate significant reductions in critical pressure, offering a scalable platform to study and tune imperfection-sensitive mechanics. Beyond shell mechanics, we offer a scalable and tunable fabrication method for patterning soft materials in applications ranging from morphable surfaces to bioinspired design.

physics.app-ph

An Empirical Study of Federated Prompt Learning for Vision Language Model

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. This paper systematically investigates the behavioral differences between language prompt learning (LPT) and vision prompt learning (VPT) under data heterogeneity challenges, including label skew and domain shift. We conduct extensive experiments to evaluate the impact of various FL and prompt configurations, such as client scale, aggregation strategies, and prompt length, to assess the robustness of Federated Prompt Learning (FPL). Furthermore, we explore strategies for enhancing prompt learning in complex scenarios where label skew and domain shift coexist, including leveraging both prompt types when computational resources allow. Our findings offer practical insights into optimizing prompt learning in federated settings, contributing to the broader deployment of VLMs in privacy-preserving environments.

cs.LG

Deployable 3D mesoscale structures through wafer fabrication, geometric frustration and bistable auxeticity

Transforming planar mesoscale devices into precise 3-D architectures is vital for next-generation flexible electronics, implants, and adaptive optics, yet wafer-based manufacturing to free-standing 3-D structures remain elusive. We fabricate polyimide architected 2-D precursors whose bistable unit cells deploy into stable 3-D mesoscale structures. Target Gaussian curvature is encoded by conformally flattening the desired mesh and locally tuning each cell so its second equilibrium matches the required scaling factor, aided by a computed library of isotropically expanding, bistable microstructures. The resulting heterogeneous tessellations uniquely morph into complex shapes. A flat disk deploys into a hemispherical dome with sub-millimeter accuracy and retains its shape after indentation. The same process yields positive- and negative-curvature geometries and tunable-focus paraboloidal mirrors whose reflected laser patterns coincide with geometric optics calculations. Our wafer-compatible, generative algorithm extends far beyond flexible substrates, enabling truly deployable, high-performance electronics and optical devices.

physics.app-ph

A general physics-constrained method for the modelling of equation's closure terms with sparse data

Accurate modeling of closure terms is a critical challenge in engineering and scientific research, particularly when data is sparse (scarse or incomplete), making widely applicable models difficult to develop. This study proposes a novel approach for constructing closure models in such challenging scenarios. We introduce a Series-Parallel Multi-Network Architecture that integrates Physics-Informed Neural Networks (PINNs) to incorporate physical constraints and heterogeneous data from multiple initial and boundary conditions, while employing dedicated subnetworks to independently model unknown closure terms, enhancing generalizability across diverse problems. These closure models are integrated into an accurate Partial Differential Equation (PDE) solver, enabling robust solutions to complex predictive simulations in engineering applications.

cs.LG

An Empirical Study of Methods for Small Object Detection from Satellite Imagery

This paper reviews object detection methods for finding small objects from remote sensing imagery and provides an empirical evaluation of four state-of-the-art methods to gain insights into method performance and technical challenges. In particular, we use car detection from urban satellite images and bee box detection from satellite images of agricultural lands as application scenarios. Drawing from the existing surveys and literature, we identify several top-performing methods for the empirical study. Public, high-resolution satellite image datasets are used in our experiments.

cs.CV

Ultra-sensitive integrated circuit sensors based on high-order nonHermitian topological physics

High-precision sensors are of fundamental importance in modern society and technology.Although numerous sensors have been developed, obtaining sensors with higher levels of sensitivity and stronger robustness has always been expected. Here, we propose theoretically and demonstrate experimentally a novel class of sensors with superior performances based on exotic properties of highorder non-Hermitian topological physics. The frequency shift induced by perturbations for these sensors can show an exponential growth with respect to the size of the device, which can well beyond the limitations of conventional sensors. The fully integrated circuit chips have been designed and fabricated in a standard 65nm complementary metal oxide semiconductor process technology. The sensitivity of systems not only less than 0.001fF has been experimentally verified, they are also robust against disorders.Our proposed ultra-sensitive integrated circuit sensors can possess a wide range of applications in various fields and show an exciting prospect for next-generation sensing technologies.

quant-ph

Multi-level mechanical modeling and computational design framework for weft knitted fabrics

This work presents a multi-level modeling and design framework for weft knitted fabrics, beginning with a volumetric finite element analysis capturing their mechanical behavior from fundamental principles. Incorporating yarn-level data, it accurately predicts stress-strain responses, reducing the need for extensive physical testing. A simplified strain energy approach homogenizes the results into three key variables, enabling rapid, accurate predictions in minutes. After validation against experiments, our framework can simulate new knit fabrics without additional tests. In real-world scenarios, fabrics often feature variations in yarn materials or patterns. The framework extends to heterogeneous fabrics, showing that transitions between distinct regions can be captured using simple mechanical analogies: springs in series and parallel. This allows heterogeneous textiles to be treated as idealized patchworks of homogeneous pieces, preserving predictive accuracy. The method is demonstrated by designing and producing a compression sleeve with uniform pressure, illustrating how the framework supports development of knits tailored to specific assistance levels and anatomical features. By combining volumetric finite element analysis, simplified model through homogenization, and controlled material transitions, this approach provides a scalable, high-fidelity path toward next-generation weft knitted fabric design.

cond-mat.soft

Personalized Quantum Federated Learning for Privacy Image Classification

Quantum federated learning has brought about the improvement of privacy image classification, while the lack of personality of the client model may contribute to the suboptimal of quantum federated learning. A personalized quantum federated learning algorithm for privacy image classification is proposed to enhance the personality of the client model in the case of an imbalanced distribution of images. First, a personalized quantum federated learning model is constructed, in which a personalized layer is set for the client model to maintain the personalized parameters. Second, a personalized quantum federated learning algorithm is introduced to secure the information exchanged between the client and server.Third, the personalized federated learning is applied to image classification on the FashionMNIST dataset, and the experimental results indicate that the personalized quantum federated learning algorithm can obtain global and local models with excellent performance, even in situations where local training samples are imbalanced. The server's accuracy is 100% with 8 clients and a distribution parameter of 100, outperforming the non-personalized model by 7%. The average client accuracy is 2.9% higher than that of the non-personalized model with 2 clients and a distribution parameter of 1. Compared to previous quantum federated learning algorithms, the proposed personalized quantum federated learning algorithm eliminates the need for additional local training while safeguarding both model and data privacy.It may facilitate broader adoption and application of quantum technologies, and pave the way for more secure, scalable, and efficient quantum distribute machine learning solutions.

quant-ph

Engineering topological states and quantum-inspired information processing using classical circuits

Based on the correspondence between circuit Laplacian and Schrodinger equation, recent investigations have shown that classical electric circuits can be used to simulate various topological physics and the Schrodinger's equation. Furthermore, a series of quantum-inspired information processing have been implemented by using classical electric circuit networks. In this review, we begin by analyzing the similarity between circuit Laplacian and lattice Hamiltonian, introducing topological physics based on classical circuits. Subsequently, we provide reviews of the research progress in quantum-inspired information processing based on the electric circuit, including discussions of topological quantum computing with classical circuits, quantum walk based on classical circuits, quantum combinational logics based on classical circuits, electric-circuit realization of fast quantum search, implementing unitary transforms and so on.

cond-mat.mes-hall

Progressive Retinal Image Registration via Global and Local Deformable Transformations

Retinal image registration plays an important role in the ophthalmological diagnosis process. Since there exist variances in viewing angles and anatomical structures across different retinal images, keypoint-based approaches become the mainstream methods for retinal image registration thanks to their robustness and low latency. These methods typically assume the retinal surfaces are planar, and adopt feature matching to obtain the homography matrix that represents the global transformation between images. Yet, such a planar hypothesis inevitably introduces registration errors since retinal surface is approximately curved. This limitation is more prominent when registering image pairs with significant differences in viewing angles. To address this problem, we propose a hybrid registration framework called HybridRetina, which progressively registers retinal images with global and local deformable transformations. For that, we use a keypoint detector and a deformation network called GAMorph to estimate the global transformation and local deformable transformation, respectively. Specifically, we integrate multi-level pixel relation knowledge to guide the training of GAMorph. Additionally, we utilize an edge attention module that includes the geometric priors of the images, ensuring the deformation field focuses more on the vascular regions of clinical interest. Experiments on two widely-used datasets, FIRE and FLoRI21, show that our proposed HybridRetina significantly outperforms some state-of-the-art methods. The code is available at https://github.com/lyp-deeplearning/awesome-retinal-registration.

cs.CV