Searcharxiv⌕ Search

arXiv subjects

Feng Yang

Publications and source records attributed to Feng Yang.

At least 73 records · Page 4Linked to original sources

MonoDETRNext: Next-Generation Accurate and Efficient Monocular 3D Object Detector

Monocular 3D object detection has vast application potential across various fields. DETR-type models have shown remarkable performance in different areas, but there is still considerable room for improvement in monocular 3D detection, especially with the existing DETR-based method, MonoDETR. After addressing the query initialization issues in MonoDETR, we explored several performance enhancement strategies, such as incorporating a more efficient encoder and utilizing a more powerful depth estimator. Ultimately, we proposed MonoDETRNext, a model that comes in two variants based on the choice of depth estimator: MonoDETRNext-E, which prioritizes speed, and MonoDETRNext-A, which focuses on accuracy. We posit that MonoDETRNext establishes a new benchmark in monocular 3D object detection and opens avenues for future research. We conducted an exhaustive evaluation demonstrating the model's superior performance against existing solutions. Notably, MonoDETRNext-A demonstrated a 3.52$\%$ improvement in the $AP_{3D}$ metric on the KITTI test benchmark over MonoDETR, while MonoDETRNext-E showed a 2.35$\%$ increase. Additionally, the computational efficiency of MonoDETRNext-E slightly exceeds that of its predecessor.

cs.CV↗

Layered semiconducting electrides in p-block metal oxides

In conventional electrides, excess electrons are localized in crystal voids to serve as anions. Most of these electrides are metallic and the metal cations are primarily from the s-block, d-block, or rare-earth elements. Here, we report a class of p-block metal-based electrides found in bilayer SnO and PbO, which are semiconducting and feature electride states in both the valence band (VB) and conduction band (CB), as referred to 2D "bipolar" electrides. These bilayers are hybrid electrides where excess electrons are localized in the interlayer region and hybridize with the orbitals of Sn atoms in the VB, exhibiting strong covalent-like interactions with neighboring metal atoms. Compared to previously studied hybrid electrides, the higher electronegativity of Sn and Pb enhances these covalent-like interactions, leading to largely enhanced semiconducting bandgap of up to 2.5 eV. Moreover, the CBM primarily arises from the overlap between metal states and interstitial charges, denoting a potential electride and forming a free-electron-like (FEL) state with small effective mass. This state offers high carrier mobilities for both electron and hole in bilayer SnO, suggesting its potential as a promising p-type semiconductor material.

cond-mat.mtrl-sci↗

Optical Diffusion Models for Image Generation

Diffusion models generate new samples by progressively decreasing the noise from the initially provided random distribution. This inference procedure generally utilizes a trained neural network numerous times to obtain the final output, creating significant latency and energy consumption on digital electronic hardware such as GPUs. In this study, we demonstrate that the propagation of a light beam through a semi-transparent medium can be programmed to implement a denoising diffusion model on image samples. This framework projects noisy image patterns through passive diffractive optical layers, which collectively only transmit the predicted noise term in the image. The optical transparent layers, which are trained with an online training approach, backpropagating the error to the analytical model of the system, are passive and kept the same across different steps of denoising. Hence this method enables high-speed image generation with minimal power consumption, benefiting from the bandwidth and energy efficiency of optical information processing.

physics.optics↗

ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling

Recognizing and disentangling visual attributes from objects is a foundation to many computer vision applications. While large vision language representations like CLIP had largely resolved the task of zero-shot object recognition, zero-shot visual attribute recognition remains a challenge because CLIP's contrastively-learned vision-language representation cannot effectively capture object-attribute dependencies. In this paper, we target this weakness and propose a sentence generation-based retrieval formulation for attribute recognition that is novel in 1) explicitly modeling a to-be-measured and retrieved object-attribute relation as a conditional probability graph, which converts the recognition problem into a dependency-sensitive language-modeling problem, and 2) applying a large pretrained Vision-Language Model (VLM) on this reformulation and naturally distilling its knowledge of image-object-attribute relations to use towards attribute recognition. Specifically, for each attribute to be recognized on an image, we measure the visual-conditioned probability of generating a short sentence encoding the attribute's relation to objects on the image. Unlike contrastive retrieval, which measures likelihood by globally aligning elements of the sentence to the image, generative retrieval is sensitive to the order and dependency of objects and attributes in the sentence. We demonstrate through experiments that generative retrieval consistently outperforms contrastive retrieval on two visual reasoning datasets, Visual Attribute in the Wild (VAW), and our newly-proposed Visual Genome Attribute Ranking (VGARank).

cs.CV↗

Purity-dependent Lorenz number, electron hydrodynamics and electron-phonon coupling in WTe$_2$

We present a study of electrical and thermal transport in Weyl semimetal WTe$_2$ down to 0.3 K. The Wiedemann-Franz law holds below 2 K and a downward deviation starts above. The deviation is more pronounced in cleaner samples, as expected in the hydrodynamic picture of electronic transport, where a fraction of electron-electron collisions conserve momentum. Phonons are the dominant heat carriers and their mean-free-path do not display a Knudsen minimum. This is presumably a consequence of weak anharmonicity, as indicated by the temperature dependence of the specific heat. Frequent momentum exchange between phonons and electrons leads to quantum oscillations of the phononic thermal conductivity. Bloch-Grüneisen picture of electron-phonon scattering breaks down at low temperature when Umklapp ph-ph collisions cease to be a sink for electronic flow of momentum. Comparison with semi-metallic Sb shows that normal ph-ph collisions are amplified by anharmonicity. In both semimetals, at cryogenic temperature, e-ph collisions degrade the phononic flow of energy but not the electronic flow of momentum.

cond-mat.mes-hall↗

Absence of BCS-BEC Crossover in FeSe0.45Te0 55 Superconductor

In iron-based superconductor Fe(Se,Te), a flat band-like feature near the Fermi level was observed around the Brillouin zone center in the superconducting state. It is under debate whether this is the evidence on the presence of the BCS-BEC crossover in the superconductor. High-resolution laser-based angle-resolved photoemission measurements are carried out on high quality single crystals of FeSe0.45Te0.55 superconductor to address the issue. By employing different polarization geometries, we have resolved and isolated the dyz band and the topological surface band, making it possible to study their superconducting behaviors separately. The dyz band alone does not form a flat band-like feature in the superconducting state and the measured dispersion can be well described by the BCS picture. We find that the flat band-like feature is formed from the combination of the dyz band and the topological surface state band in the superconducting state. These results reveal the origin of the flat band-like feature and rule out the presence of BCS-BEC crossover in Fe(Se,Te) superconductor.

cond-mat.supr-con↗

Fluid-Antenna Enhanced ISAC: Joint Antenna Positioning and Dual-Functional Beamforming Design under Perfect and Imperfect CSI

Integrated sensing and communication (ISAC) emerges as an essential technique for overcoming spectrum congestion. However, the performance of traditional ISAC systems with fixed-position-antennas (FPA) is limited due to insufficient spatial degree of freedom (DoF) exploration. Recently, fluid antenna (FA) with reconfigurable antenna position is developed to enhance the sensing and communication performance by reshaping the channel. This paper investigates an FA-enhanced ISAC system where a base station is equipped with multiple FAs to communicate with multiple single-antenna users and with FPAs to sense a point target. In this paper, we consider both perfect and imperfect channel state information (CSI) of the communication channel and sensing channel. In two cases, we focus on the maximization of the sensing signal-to-noise (SNR) by optimizing the positions of FAs and the dual-functional beamforming under the constraints of the FA moving region, the minimum FA distance and the minimum signal-to-interference-plus-noise (SINR) per user. Specifically, for the ideal case of perfect CSI, an iterative alternating optimization (AO) algorithm is proposed to tackle the formulated problem where the dual-functional beamforming and the FA positions are obtained via semidefinite relaxation (SDR) and successive convex approximation (SCA) techniques. Then, for the imperfect CSI case, we propose an AO-based iterative algorithm where $\mathcal{S}-$Procedure and SCA are applied to obtain the dual-functional beamforming and the FA positions. Furthermore, we analytically and numerically prove the convergence of the proposed algorithms. Numerical results demonstrate the notable gains of the proposed algorithms in the respective cases.

eess.SP↗

Negligible Normal Fluid in Superconducting State of Heavily Overdoped Bi$_2$Sr$_2$CaCu$_2$O$_{8+δ}$ Detected by Ultra-Low Temperature Angle-Resolved Photoemission Spectroscopy

In high temperature cuprate superconductors, it was found that in the overdoped region the superfluid density decreases with the increase of hole doping. One natural question is whether there exists normal fluid in the superconducting state in the overdoped region. In this paper, we have carried out high-resolution ultra-low temperature laser-based angle-resolved photoemission measurements on a heavily overdoped Bi2212 sample with a $T_{\mathrm{c}}$ of 48 K. We find that this heavily overdoped Bi2212 remains in the strong coupling regime with $2 \mathitΔ_0 / k_{\mathrm{B}} T_{\mathrm{c}}=5.8$. The single-particle scattering rate is very small along the nodal direction ($\sim$5 meV) and increases as the momentum moves from the nodal to the antinodal regions. A hard superconducting gap opening is observed near the antinodal region with the spectral weight at the Fermi level fully suppressed to zero. The normal fluid is found to be negligibly small in the superconducting state of this heavily overdoped Bi2212. These results provide key information to understand the high $T_\mathrm{c}$ mechanism in the cuprate superconductors.

cond-mat.supr-con↗

Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation

Recent works have demonstrated that using reinforcement learning (RL) with multiple quality rewards can improve the quality of generated images in text-to-image (T2I) generation. However, manually adjusting reward weights poses challenges and may cause over-optimization in certain metrics. To solve this, we propose Parrot, which addresses the issue through multi-objective optimization and introduces an effective multi-reward optimization strategy to approximate Pareto optimal. Utilizing batch-wise Pareto optimal selection, Parrot automatically identifies the optimal trade-off among different rewards. We use the novel multi-reward optimization algorithm to jointly optimize the T2I model and a prompt expansion network, resulting in significant improvement of image quality and also allow to control the trade-off of different rewards using a reward related prompt during inference. Furthermore, we introduce original prompt-centered guidance at inference time, ensuring fidelity to user input after prompt expansion. Extensive experiments and a user study validate the superiority of Parrot over several baselines across various quality criteria, including aesthetics, human preference, text-image alignment, and image sentiment.

cs.CV↗

WateRF: Robust Watermarks in Radiance Fields for Protection of Copyrights

The advances in the Neural Radiance Fields (NeRF) research offer extensive applications in diverse domains, but protecting their copyrights has not yet been researched in depth. Recently, NeRF watermarking has been considered one of the pivotal solutions for safely deploying NeRF-based 3D representations. However, existing methods are designed to apply only to implicit or explicit NeRF representations. In this work, we introduce an innovative watermarking method that can be employed in both representations of NeRF. This is achieved by fine-tuning NeRF to embed binary messages in the rendering process. In detail, we propose utilizing the discrete wavelet transform in the NeRF space for watermarking. Furthermore, we adopt a deferred back-propagation technique and introduce a combination with the patch-wise loss to improve rendering quality and bit accuracy with minimum trade-offs. We evaluate our method in three different aspects: capacity, invisibility, and robustness of the embedded watermarks in the 2D-rendered images. Our method achieves state-of-the-art performance with faster training speed over the compared state-of-the-art methods.

cs.CV↗

Fluid-Antenna Enhanced Integrated Sensing and Communication: Joint Antenna Positioning and Beamforming Design

This paper investigates a fluid antenna (FA) enhanced integrated sensing and communication (ISAC) system consisting of a base station (BS), multiple single-antenna communication users, and one point target, where the BS is equipped with FAs to enhance both the communication and sensing performance. First, we formulate a problem that maximizes the radar signal-to-noise ratio (SNR) by jointly optimizing the FAs' positions and transmit beamforming matrix. Then, to tackle this highly non-convex problem, we present efficient algorithms by using alternating optimization (AO), successive convex approximation (SCA), and semi-definite relaxation (SDR). Numerical results demonstrate the convergence behavior and effectiveness of the proposed algorithm.

eess.SP↗

Integrated Modeling, Verification, and Code Generation for Unmanned Aerial Systems

Unmanned Aerial Systems (UAS) are currently widely used in safety-critical fields such as industrial production, military operations, and disaster relief. Due to the diversity and complexity of application scenarios, UAS have become increasingly intricate. The challenge of designing and implementing highly reliable UAS while effectively controlling development costs and enhancing efficiency is a pressing issue faced by both academia and industry. Addressing this challenge, this paper aims to investigate an integrated approach to modeling, verification, and code generation for UAS. The paper begins by utilizing Architecture Analysis and Design Language (AADL) to model the UAS, proposing a set of generic UAS models. Based on these models, formal specifications are written to describe the system's safety properties and functions. Finally, the paper introduces a method for generating flight controller code for UAS based on the verified models. Experiments conducted with the proposed method demonstrate its effectiveness in identifying potential vulnerabilities in the UAS during the early design phase and in generating viable flight controller code from the verified models. This approach can enhance the efficiency of designing and verifying high-reliability UAS.

cs.SE↗

A Survey of Deep Learning Based Radar and Vision Fusion for 3D Object Detection in Autonomous Driving

With the rapid advancement of autonomous driving technology, there is a growing need for enhanced safety and efficiency in the automatic environmental perception of vehicles during their operation. In modern vehicle setups, cameras and mmWave radar (radar), being the most extensively employed sensors, demonstrate complementary characteristics, inherently rendering them conducive to fusion and facilitating the achievement of both robust performance and cost-effectiveness. This paper focuses on a comprehensive survey of radar-vision (RV) fusion based on deep learning methods for 3D object detection in autonomous driving. We offer a comprehensive overview of each RV fusion category, specifically those employing region of interest (ROI) fusion and end-to-end fusion strategies. As the most promising fusion strategy at present, we provide a deeper classification of end-to-end fusion methods, including those 3D bounding box prediction based and BEV based approaches. Moreover, aligning with recent advancements, we delineate the latest information on 4D radar and its cutting-edge applications in autonomous vehicles (AVs). Finally, we present the possible future trends of RV fusion and summarize this paper.

cs.CV↗

Orbital-Dependent Electron Correlation in Double-Layer Nickelate La3Ni2O7

The latest discovery of high temperature superconductivity near 80K in La3Ni2O7 under high pressure has attracted much attention. Many proposals are put forth to understand the origin of superconductivity.The determination of electronic structures is a prerequisite to establish theories to understand superconductivity in nickelates but is still lacking. Here we report our direct measurement of the electronic structures of La3Ni2O7 by high-resolution angle resolved photoemission spectroscopy. The Fermi surface and band structures of La3Ni2O7 are observed and compared with the band structure calculations. Strong electron correlations are revealed which are orbital- and momentum dependent. A flat band is formed from the Ni-3dz2 orbitals around the zone corner which is ~50meV below the Fermi level and exhibits the strongest electron correlation. In many theoretical proposals, this band is expected to play the dominant role in generating superconductivity in La3Ni2O7. Our observations provide key experimental information to understand the electronic structure and origin of high temperature superconductivity in La3Ni2O7.

cond-mat.supr-con↗

An Agile Formal Specification Language Design Based on K Framework

Formal Methods (FMs) are currently essential for verifying the safety and reliability of software systems. However, the specification writing in formal methods tends to be complex and challenging to learn, requiring familiarity with various intricate formal specification languages and verification technologies. In response to the increasing complexity of software frameworks, existing specification writing methods fall short in meeting agility requirements. To address this, this paper introduces an Agile Formal Specification Language (ASL). The ASL is defined based on the K Framework and YAML Ain't Markup Language (YAML). The design of ASL incorporates agile design principles, making the writing of formal specifications simpler, more efficient, and scalable. Additionally, a specification translation algorithm is developed, capable of converting ASL into K formal specification language that can be executed for verification. Experimental evaluations demonstrate that the proposed method significantly reduces the code size needed for specification writing, enhancing agility in formal specification writing.

cs.SE↗

Rich Human Feedback for Text-to-Image Generation

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However, many generated images still suffer from issues such as artifacts/implausibility, misalignment with text descriptions, and low aesthetic quality. Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) for large language models, prior works collected human-provided scores as feedback on generated images and trained a reward model to improve the T2I generation. In this paper, we enrich the feedback signal by (i) marking image regions that are implausible or misaligned with the text, and (ii) annotating which words in the text prompt are misrepresented or missing on the image. We collect such rich human feedback on 18K generated images (RichHF-18K) and train a multimodal transformer to predict the rich feedback automatically. We show that the predicted rich human feedback can be leveraged to improve image generation, for example, by selecting high-quality training data to finetune and improve the generative models, or by creating masks with predicted heatmaps to inpaint the problematic regions. Notably, the improvements generalize to models (Muse) beyond those used to generate the images on which human feedback data were collected (Stable Diffusion variants). The RichHF-18K data set will be released in our GitHub repository: https://github.com/google-research/google-research/tree/master/richhf_18k.

cs.CV↗

Intrinsic conductance of ferroelectric domain walls

Ferroelectric domain walls hold great promise for innovative applications in ferroelectric devices. However, the underlying mechanisms behind the observed giant conductance of charged domain walls remain poorly understood. Using a first-principles approach that incorporates Boltzmann transport theory and the relaxation time approximation, we determine the carrier concentration, mobility, and conductivity of domain walls with head-to-head and tail-to-tail polarization orientations. Our systematic exploration reveals that the accumulation of carriers, particularly their concentration, plays a dominant role in the domain wall conductance mechanism. However, the observed conductance differences between head-to-head and tail-to-tail domain walls are primarily due to differences in carrier mobility. The width of the domain wall is a key factor determining the device scale. Our calculated domain wall width is significantly smaller than previously reported values. This method, not limited to a certain ferroelectric material, can be used for the optimization and application development of various domain wall materials and devices.

cond-mat.mtrl-sci↗