SearcharxivSearch

arXiv subjects

Peng

Publications and source records attributed to Peng.

7 recordsLinked to original sources

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning

Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal reasoning can lead to instability, due to the exploitation of language priors, the neglect of visual evidence, and the generation of reasoning traces that are fluent yet not visually grounded. The question arises: Can initially steer the policy toward visually faithful reasoning regime before applying reinforcement learning? To this end, we propose a Faithful Warm-Start (FWS) strategy that first curates samples with explicit vision-language causal relationships from six general VQA benchmarks to construct the FaithfulQA dataset, where each of the image-question pairs gains a certain degree of visual observations, question requirements, commonsense knowledge, domain knowledge, and the final answer. Subsequently, a VLM-based judge is employed to further purify the dataset, ensuring strong causal consistency and visual faithfulness. This warm-start stage equips the model with the capability to understand causally grounded vision-language patterns before subsequent RL optimization under sparse answer-level rewards. Experimental results show that such faithful supervision improves answer accuracy, stabilizes RL training, and reduces visually unsupported reasoning.

cs.AI

CONTRACTFIX: A Framework for Automatically Fixing Vulnerabilities in Smart Contracts

The increased adoption of smart contracts in many industries has made them an attractive target for cybercriminals, leading to millions of dollars in losses. Thus, deploying smart contracts with detected vulnerabilities (known to developers) are not acceptable, and fixing all the detected vulnerabilities is needed, which incurs high manual labor cost without effective tool support. To fill this need, in this paper, we propose ContractFix, a novel framework that automatically generates security patches for vulnerable smart contracts. ContractFix is a general framework that can incorporate different fix patterns for different types of vulnerabilities. Users can use it as a security fix-it tool that automatically applies patches and verifies the patched contracts before deploying the contracts. To address the unique challenges in fixing smart contract vulnerabilities, given an input smart contract, \tool conducts our proposed ensemble identification based on multiple static verification tools to identify vulnerabilities that are amenable for automatic fix. Then, ContractFix generates patches using template-based fix patterns and conducts program analysis (program dependency computation and pointer analysis) for smart contracts to accurately infer and populate the parameter values for the fix patterns. Finally, ContractFix performs static verification that guarantees the patched contract is free of vulnerabilities. Our evaluations on $144$ real vulnerable contracts demonstrate that \tool can successfully fix $94\%$ of the detected vulnerabilities ($565$ out of $601$) and preserve the expected behaviors of the smart contracts.

cs.SE

High capacity topological coding based on nested vortex knots and links

Optical knots and links have attracted great attention because of their exotic topological characteristics. Recent investigations have shown that the information encoding based on optical knots could possess robust features against external perturbations. However, as a superior coding scheme, it is also necessary to achieve a high capacity, which is hard to be fulfilled by existing knot-carriers owing to the limit number of associated topological invariants. Thus, how to realize the knot-based information coding with a high capacity is a key problem to be solved. Here, we create a type of nested vortex knot, and show that it can be used to fulfill the robust information coding with a high capacity assisted by a large number of intrinsic topological invariants. In experiments, we design and fabricate metasurface holograms to generate light fields sustaining different kinds of nested vortex links. Furthermore, we verify the feasibility of the high-capacity coding scheme based on those topological optical knots. Our work opens another way to realize the robust and high capacity optical coding, which may have useful impacts on the field of information transfer and storage.

physics.optics

Time-multiplexed Neural Holography: A flexible framework for holographic near-eye displays with fast heavily-quantized spatial light modulators

Holographic near-eye displays offer unprecedented capabilities for virtual and augmented reality systems, including perceptually important focus cues. Although artificial intelligence--driven algorithms for computer-generated holography (CGH) have recently made much progress in improving the image quality and synthesis efficiency of holograms, these algorithms are not directly applicable to emerging phase-only spatial light modulators (SLM) that are extremely fast but offer phase control with very limited precision. The speed of these SLMs offers time multiplexing capabilities, essentially enabling partially-coherent holographic display modes. Here we report advances in camera-calibrated wave propagation models for these types of holographic near-eye displays and we develop a CGH framework that robustly optimizes the heavily quantized phase patterns of fast SLMs. Our framework is flexible in supporting runtime supervision with different types of content, including 2D and 2.5D RGBD images, 3D focal stacks, and 4D light fields. Using our framework, we demonstrate state-of-the-art results for all of these scenarios in simulation and experiment.

cs.GR

Color Contrast Enhanced Rendering for Optical See-through Head-mounted Displays

Most commercially available optical see-through head-mounted displays (OST-HMDs) utilize optical combiners to simultaneously visualize the physical background and virtual objects. The displayed images perceived by users are a blend of rendered pixels and background colors. Enabling high fidelity color perception in mixed reality (MR) scenarios using OST-HMDs is an important but challenging task. We propose a real-time rendering scheme to enhance the color contrast between virtual objects and the surrounding background for OST-HMDs. Inspired by the discovery of color perception in psychophysics, we first formulate the color contrast enhancement as a constrained optimization problem. We then design an end-to-end algorithm to search the optimal complementary shift in both chromaticity and luminance of the displayed color. This aims at enhancing the contrast between virtual objects and the real background as well as keeping the consistency with the original color. We assess the performance of our approach using a simulated OST-HMD environment and an off-the-shelf OST-HMD. Experimental results from objective evaluations and subjective user studies demonstrate that the proposed approach makes rendered virtual objects more distinguishable from the surrounding background, thereby bringing a better visual experience.

cs.GR

Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data

In many applications, one works with neural network models trained by someone else. For such pretrained models, one may not have access to training data or test data. Moreover, one may not know details about the model, e.g., the specifics of the training data, the loss function, the hyperparameter values, etc. Given one or many pretrained models, it is a challenge to say anything about the expected performance or quality of the models. Here, we address this challenge by providing a detailed meta-analysis of hundreds of publicly-available pretrained models. We examine norm based capacity control metrics as well as power law based metrics from the recently-developed Theory of Heavy-Tailed Self Regularization. We find that norm based metrics correlate well with reported test accuracies for well-trained models, but that they often cannot distinguish well-trained versus poorly-trained models. We also find that power law based metrics can do much better -- quantitatively better at discriminating among series of well-trained models with a given architecture; and qualitatively better at discriminating well-trained versus poorly-trained models. These methods can be used to identify when a pretrained neural network has problems that cannot be detected simply by examining training/test accuracies.

cs.LG

SpArcFiRe: morphological selection effects due to reduced visibility of tightly winding arms in distant spiral galaxies

The Galaxy Zoo has provided morphological data on many galaxies. Several biases have been identified in the Galaxy Zoo data. Here we report on a newly discovered selection effect: astronomers interested in studying spiral galaxies may select a set of spiral galaxies based upon a threshold in spirality (the fraction of Galaxy Zoo humans who report seeing spiral structure). SpArcFiRe is an automated tool that decomposes a spiral galaxy into its constituent spiral arms, providing objective, quantitative data on their structure. SpArcFiRe measures the pitch angle of spiral arms. We have observed that when selecting a set of spiral galaxies based on a threshold on spirality, the pitch angle of spiral arms appear increase with redshift. We hypothesize that this is a selection effect: tightly-wound spiral arms become less visible as images degrade with increasing redshift, leading to fewer such galaxies being included in the sample at higher redshifts. We corroborate this hypothesis by artificially degrading images of nearby galaxies, then using a machine learning algorithm trained on Galaxy Zoo data to provide a spirality for each artificially degraded image. It correctly predicts that spirality decreases as image quality degrades. Thus, the mean pitch angle of those galaxies remaining above the spirality threshold is higher than those eliminated by the selection effect. This demonstrates that users who select samples of galaxies using a threshold of Galaxy Zoo votes must carefully consider the possibility of selection effects on morphological measures, even if the measure itself is believed to be objective and unbiased. Finally, we also perform an empirical sensitivity analysis to demonstrate that SpArcFiRe's output changes in a smooth and predictable fashion to changes in its internal algorithmic parameters.

astro-ph.GA