SearcharxivSearch

arXiv subjects

Jing Lin

Publications and source records attributed to Jing Lin.

At least 55 records · Page 3Linked to original sources

BinaryHPE: 3D Human Pose and Shape Estimation via Binarization

3D human pose and shape estimation (HPE) aims to reconstruct the 3D human body, face, and hands from a single image. Although powerful deep learning models have achieved accurate estimation in this task, they require enormous memory and computational resources. Consequently, these methods can hardly be deployed on resource-limited edge devices. In this work, we propose BinaryHPE, a novel binarization method designed to estimate the 3D human body, face, and hands parameters efficiently. Specifically, we propose a novel binary backbone called Binarized Dual Residual Network (BiDRN), designed to retain as much full-precision information as possible. Furthermore, we propose the Binarized BoxNet, an efficient sub-network for predicting face and hands bounding boxes, which further reduces model redundancy. Comprehensive quantitative and qualitative experiments demonstrate the effectiveness of BinaryHPE, which has a significant improvement over state-of-the-art binarization algorithms. Moreover, our BinaryHPE achieves comparable performance with the full-precision method Hand4Whole while using only 22.1% parameters and 14.8% operations. We will release all the code and pretrained models.

cs.CV

Algebraic skin effect in two-dimensional non-Hermitian metamaterials

Metamaterials have unlocked unprecedented control over light by leveraging novel mechanisms to expand their functionality. Non-Hermitian physics further enhances the tunability of non-Hermitian metamaterials (NHMs) through phenomena such as the non-Hermitian skin effect (NHSE), enabling applications like directional amplification. The higher-dimensional NHSE manifests unique effects, including the algebraic skin effect (ASE), which features power-law decay instead of exponential localization, allowing for quasi-long-range interactions. In this work, we establish apparent criteria for achieving ASE in two-dimensional reciprocal NHMs with anisotropic and complex dielectric tensors. By numerically and theoretically demonstrating ASE through mismatched optical axes and geometric structures, we reveal that ASE is governed by a generalized Fermi surface whose dimensionality exceeds that of the Fermi surface. We further propose and validate a realistic photonic crystal design for ASE, which is experimentally accessible. Our recipe for ASE provides a versatile pathway for broader generalizations, including three-dimensional structures, synthetic dimensions, and other classical wave systems, paving the way for advancements in non-Hermitian photonics.

physics.optics

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset

In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset. Existing motion datasets predominantly capture body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions, and are typically limited to lab settings with manually labeled text descriptions, thereby restricting their scalability. To address this issue, we develop a scalable annotation pipeline that can automatically capture 3D whole-body human motion and comprehensive textural labels from RGB videos and build the Motion-X dataset comprising 81.1K text-motion pairs. Furthermore, we extend Motion-X into Motion-X++ by improving the annotation pipeline, introducing more data modalities, and scaling up the data quantities. Motion-X++ provides 19.5M 3D whole-body pose annotations covering 120.5K motion sequences from massive scenes, 80.8K RGB videos, 45.3K audios, 19.5M frame-level whole-body pose descriptions, and 120.5K sequence-level semantic labels. Comprehensive experiments validate the accuracy of our annotation pipeline and highlight Motion-X++'s significant benefits for generating expressive, precise, and natural motion with paired multimodal labels supporting several downstream tasks, including text-driven whole-body motion generation,audio-driven motion generation, 3D whole-body human mesh recovery, and 2D whole-body keypoints estimation, etc.

cs.CV

Unveiling Non-Hermitian Spectral Topology in Hyperbolic Lattices with Non-Abelian Translation Symmetry

The hyperbolic lattice (HBL) has emerged as a compelling platform for exploring matter in non-Euclidean space. Among its notable features, the breakdown of the conventional Bloch theorem stands out, prompting a reexamination of band theory, with the determination of spectra for non-Hermitian systems being a prominent example. Here, we develop an approach to determining the spectra under open boundary conditions (OBCs), one of the foundations in non-Hermitian lattices, from the reciprocal space of HBLs. By introducing supercells to encompass states that are allowed by non-Abelian translational groups, we perform analytic continuation and base on the point gap topology to acquire uniform spectra, the universal OBC spectral range. Applying this method to a single-band nonreciprocal model and a reciprocal non-Abelian semimetal model, we reveal higher-dimensional skin effects and topological phase transitions, respectively, demonstrating the feasibility of our method in predicting spectral topology and investigating non-Hermitian physics in HBLs.

cond-mat.mes-hall

Zeolitic Imidazolate Framework-8 offers an anti-inflammatory and antifungal method in the treatment of Aspergillus fungus keratitis in vitro and in vivo

Background: Fungal keratitis is a serious blinding eye disease. Traditional drugs used to treat fungal keratitis commonly have the disadvantages of low bioavailability, poor dispersion, and limited permeability. Purpose: To develop a new method for the treatment of fungal keratitis with improved bioavailability, dispersion, and permeability. Purpose: To develop a new method for the treatment of fungal keratitis with improved bioavailability, dispersion, and permeability. Methods: Zeolitic Imidazolate Framework-8 (ZIF-8) was formed by zinc ions and 2-methylimidazole linked by coordination bonds and characterized by Scanning electron microscopy (SEM), X-ray diffraction (XRD), and Zeta potential. The safety of ZIF-8 on HCECs and RAW 264.7 cells was detected by Cell Counting Kit-8 (CCK-8). The anti-inflammatory effects of ZIF-8 on RAW 246.7 cells were evaluated by Quantitative Real-Time PCR Experiments (qPCR) and Enzyme-linked immunosorbent assay (ELISA). Clinical score, Colony-Forming Units (CFU). In vivo, treatment with ZIF-8 reduced corneal fungal load and mitigated neutrophil infiltration in fungal keratitis, which effectively reduced the severity of keratitis in mice and alleviated the infiltration of inflammatory factors in the mouse cornea. In addition, ZIF-8 reduces the inflammatory response by downregulating the expression of pro-inflammatory cytokines TNF-α, IL-6, and IL-1\b{eta} after Aspergillus fumigatus infection in vivo and in vitro. Conclusion: ZIF-8 has a significant anti-inflammatory and antifungal effect, which provides a new solution for the treatment of fungal keratitis.

q-bio.TO

ZIF-90 treats fungal keratitis by promoting macrophage apoptosis and inhibiting inflammatory response

Fungal keratitis is a severe vision-threatening corneal infection with a prognosis influenced by fungal virulence and the host's immune defense mechanisms. The immune system, through its regulation of the inflammatory response, ensures cells and tissues can effectively activate defense mechanisms in response to infection and injury. However, there is still a lack of effective drugs that attenuate fungal virulence while relieving the inflammatory response caused by fungal keratitis. Therefore, finding effective treatments to solve these problems is particularly important. We synthesized ZIF-90 by water-based synthesis and characterized by SEM, XRD etc. In vitro experiments included CCK-8 and ELISA. These evaluations verified the disruptive effects of ZIF-90 on Aspergillus. fumigatus spore adhesion, morphology, cell membrane, and the effect of ZIF-90 on apoptosis. In addition, to investigate whether the metal-ligand zinc and the organic ligand imidazole act as essential factors in ZIF-90, we investigated the in vitro antimicrobial and anti-inflammatory effects of ZIF-8, ZIF-67, and MOF-74 (Zn) by MIC and ELISA experiments. ZIF-90 has therapeutic effects on fungal keratitis, which could break the protective organelles of Aspergillus. fumigatus, such as the cell wall. In addition, ZIF-90 can avoid excessive inflammatory response by promoting apoptosis of inflammatory cells. The results demonstrated that both zinc ions and imidazole possessed antimicrobial and anti-inflammatory effects. In addition, ZIF-90 exhibited better biocompatibility compared to ZIF-8, ZIF-67, and MOF-74 (Zn). ZIF-90 has anti-inflammatory and antifungal effects and preferable biocompatibility, and has great potential for the treatment of fungal keratitis.

q-bio.SC

Enhanced Digital Twin for Human-Centric and Integrated Lighting Asset Management in Public Libraries: From Corrective to Predictive Maintenance

Lighting asset management in public libraries has traditionally been reactive, focusing on corrective maintenance, addressing issues only when failures occur. Although standards now encourage preventive measures, such as incorporating a maintenance factor, the broader goal of human centric, sustainable lighting systems requires a shift toward predictive maintenance strategies. This study introduces an enhanced digital twin model designed for the proactive management of lighting assets in public libraries. By integrating descriptive, diagnostic, predictive, and prescriptive analytics, the model enables a comprehensive, multilevel view of asset health. The proposed framework supports both preventive and predictive maintenance strategies, allowing for early detection of issues and the timely resolution of potential failures. In addition to the specific application for lighting systems, the design is adaptable for other building assets, providing a scalable solution for integrated asset management in various public spaces.

cs.HC

In situ mixer calibration for superconducting quantum circuits

Mixers play a crucial role in superconducting quantum computing, primarily by facilitating frequency conversion of signals to enable precise control and readout of quantum states. However, imperfections, particularly carrier leakage and unwanted sideband signal, can significantly compromise control fidelity. To mitigate these defects, regular and precise mixer calibrations are indispensable, yet they pose a formidable challenge in large-scale quantum control. Here, we introduce an in situ calibration technique and outcome-focused mixer calibration scheme using superconducting qubits. Our method leverages the qubit's response to imperfect signals, allowing for calibration without modifying the wiring configuration. We experimentally validate the efficacy of this technique by benchmarking single-qubit gate fidelity and qubit coherence time.

quant-ph

RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

The memory and computational demands of Key-Value (KV) cache present significant challenges for deploying long-context language models. Previous approaches attempt to mitigate this issue by selectively dropping tokens, which irreversibly erases critical information that might be needed for future queries. In this paper, we propose a novel compression technique for KV cache that preserves all token information. Our investigation reveals that: i) Most attention heads primarily focus on the local context; ii) Only a few heads, denoted as retrieval heads, can essentially pay attention to all input tokens. These key observations motivate us to use separate caching strategy for attention heads. Therefore, we propose RazorAttention, a training-free KV cache compression algorithm, which maintains a full cache for these crucial retrieval heads and discards the remote tokens in non-retrieval heads. Furthermore, we introduce a novel mechanism involving a "compensation token" to further recover the information in the dropped tokens. Extensive evaluations across a diverse set of large language models (LLMs) demonstrate that RazorAttention achieves a reduction in KV cache size by over 70% without noticeable impacts on performance. Additionally, RazorAttention is compatible with FlashAttention, rendering it an efficient and plug-and-play solution that enhances LLM inference efficiency without overhead or retraining of the original model.

cs.LG

A Novel Context driven Critical Integrative Levels (CIL) Approach: Advancing Human-Centric and Integrative Lighting Asset Management in Public Libraries with Practical Thresholds

This paper proposes the context driven Critical Integrative Levels (CIL), a novel approach to lighting asset management in public libraries that aligns with the transformative vision of human-centric and integrative lighting. This approach encompasses not only the visual aspects of lighting performance but also prioritizes the physiological and psychological well-being of library users. Incorporating a newly defined metric, Mean Time of Exposure (MTOE), the approach quantifies user-light interaction, enabling tailored lighting strategies that respond to diverse activities and needs in library spaces. Case studies demonstrate how the CIL matrix can be practically applied, offering significant improvements over conventional methods by focusing on optimized user experiences from both visual impacts and non-visual effects.

cs.HC

ChatPose: Chatting about 3D Human Pose

We introduce ChatPose, a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from a single image or a brief description, a process that intertwines image interpretation, world knowledge, and an understanding of body language. Traditional human pose estimation and generation methods often operate in isolation, lacking semantic understanding and reasoning abilities. ChatPose addresses these limitations by embedding SMPL poses as distinct signal tokens within a multimodal LLM, enabling the direct generation of 3D body poses from both textual and visual inputs. Leveraging the powerful capabilities of multimodal LLMs, ChatPose unifies classical 3D human pose and generation tasks while offering user interactions. Additionally, ChatPose empowers LLMs to apply their extensive world knowledge in reasoning about human poses, leading to two advanced tasks: speculative pose generation and reasoning about pose estimation. These tasks involve reasoning about humans to generate 3D poses from subtle text queries, possibly accompanied by images. We establish benchmarks for these tasks, moving beyond traditional 3D pose generation and estimation methods. Our results show that ChatPose outperforms existing multimodal LLMs and task-specific methods on these newly proposed tasks. Furthermore, ChatPose's ability to understand and generate 3D human poses based on complex reasoning opens new directions in human pose analysis.

cs.CV

NTIRE 2024 Challenge on Low Light Image Enhancement: Methods and Results

This paper reviews the NTIRE 2024 low light image enhancement challenge, highlighting the proposed solutions and results. The aim of this challenge is to discover an effective network design or solution capable of generating brighter, clearer, and visually appealing results when dealing with a variety of conditions, including ultra-high resolution (4K and beyond), non-uniform illumination, backlighting, extreme darkness, and night scenes. A notable total of 428 participants registered for the challenge, with 22 teams ultimately making valid submissions. This paper meticulously evaluates the state-of-the-art advancements in enhancing low-light images, reflecting the significant progress and creativity in this field.

cs.CV

Human-Centric and Integrative Lighting Asset Management in Public Libraries: Qualitative Insights and Challenges from a Swedish Field Study

Traditional lighting source reliability evaluations, often covering just half of a lamp's volume, can misrepresent real-world performance. To overcome these limitations,adopting advanced asset management strategies for a more holistic evaluation is crucial. This paper investigates human-centric and integrative lighting asset management in Swedish public libraries. Through field observations, interviews, and gap analysis, the study highlights a disparity between current lighting conditions and stakeholder expectations, with issues like eye strain suggesting significant improvement potential. We propose a shift towards more dynamic lighting asset management and reliability evaluations, emphasizing continuous enhancement and comprehensive training in human-centric and integrative lighting principles.

stat.OT

Identifiability Study of Lithium-Ion Battery Capacity Fade Using Degradation Mode Sensitivity for a Minimally and Intuitively Parametrized Electrode-Specific Cell Open-Circuit Voltage Model

When two electrode open-circuit potentials form a full-cell OCV (open-circuit voltage) model, cell-level SOH (state of health) parameters related to LLI (loss of lithium inventory) and LAM (loss of active materials) naturally appear. Such models have been used to interpret experimental OCV measurements and infer these SOH parameters associated with capacity fade. In this work, we first re-parametrize a popular OCV model formulation by the N/P (negative-to-positive) ratio and Li/P (lithium-to-positive) ratio, which have more symmetric and intuitive physical meaning, and are also pristine-condition-agnostic and cutoff-voltage-independent. We then study the modal identifiability of capacity fade by mathematically deriving the gradients of electrode slippage and cell OCV with respect to these SOH parameters, where the electrode differential voltage fractions, which characterize each electrode's relative contribution to the OCV slope, play a key role in passing the influence of a fixed cutoff voltage to the parameter sensitivity. The sensitivity gradients of the total capacity also reveal four characteristic regimes regarding how much lithium inventory and active materials are limiting the apparent capacity. We show the usefulness of these sensitivity gradients with an application regarding degradation mode identifiability from OCV measurements at different SOC (state of charge) windows.

physics.chem-ph

DPoser: Diffusion Model as Robust 3D Human Pose Prior

This work targets to construct a robust human pose prior. However, it remains a persistent challenge due to biomechanical constraints and diverse human movements. Traditional priors like VAEs and NDFs often exhibit shortcomings in realism and generalization, notably with unseen noisy poses. To address these issues, we introduce DPoser, a robust and versatile human pose prior built upon diffusion models. DPoser regards various pose-centric tasks as inverse problems and employs variational diffusion sampling for efficient solving. Accordingly, designed with optimization frameworks, DPoser seamlessly benefits human mesh recovery, pose generation, pose completion, and motion denoising tasks. Furthermore, due to the disparity between the articulated poses and structured images, we propose truncated timestep scheduling to enhance the effectiveness of DPoser. Our approach demonstrates considerable enhancements over common uniform scheduling used in image domains, boasting improvements of 5.4%, 17.2%, and 3.8% across human mesh recovery, pose completion, and motion denoising, respectively. Comprehensive experiments demonstrate the superiority of DPoser over existing state-of-the-art pose priors across multiple tasks.

cs.CV

Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset

In this paper, we present Motion-X, a large-scale 3D expressive whole-body motion dataset. Existing motion datasets predominantly contain body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions. Moreover, they are primarily collected from limited laboratory scenes with textual descriptions manually labeled, which greatly limits their scalability. To overcome these limitations, we develop a whole-body motion and text annotation pipeline, which can automatically annotate motion from either single- or multi-view videos and provide comprehensive semantic labels for each video and fine-grained whole-body pose descriptions for each frame. This pipeline is of high precision, cost-effective, and scalable for further research. Based on it, we construct Motion-X, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes. Besides, Motion-X provides 15.6M frame-level whole-body pose descriptions and 81.1K sequence-level semantic labels. Comprehensive experiments demonstrate the accuracy of the annotation pipeline and the significant benefit of Motion-X in enhancing expressive, diverse, and natural motion generation, as well as 3D whole-body human mesh recovery.

cs.CV

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion allows for controllable image editing, while the integration of OSX facilitates promptable 3D human motion analysis. Grounded SAM also shows superior performance on open-vocabulary benchmarks, achieving 48.7 mean AP on SegInW (Segmentation in the wild) zero-shot benchmark with the combination of Grounding DINO-Base and SAM-Huge models.

cs.CV

PhysHOI: Physics-Based Imitation of Dynamic Human-Object Interaction

Humans interact with objects all the time. Enabling a humanoid to learn human-object interaction (HOI) is a key step for future smart animation and intelligent robotics systems. However, recent progress in physics-based HOI requires carefully designed task-specific rewards, making the system unscalable and labor-intensive. This work focuses on dynamic HOI imitation: teaching humanoid dynamic interaction skills through imitating kinematic HOI demonstrations. It is quite challenging because of the complexity of the interaction between body parts and objects and the lack of dynamic HOI data. To handle the above issues, we present PhysHOI, the first physics-based whole-body HOI imitation approach without task-specific reward designs. Except for the kinematic HOI representations of humans and objects, we introduce the contact graph to model the contact relations between body parts and objects explicitly. A contact graph reward is also designed, which proved to be critical for precise HOI imitation. Based on the key designs, PhysHOI can imitate diverse HOI tasks simply yet effectively without prior knowledge. To make up for the lack of dynamic HOI scenarios in this area, we introduce the BallPlay dataset that contains eight whole-body basketball skills. We validate PhysHOI on diverse HOI tasks, including whole-body grasping and basketball skills.

cs.CV