SearcharxivSearch

arXiv subjects

Bo Ma

Publications and source records attributed to Bo Ma.

At least 19 recordsLinked to original sources

Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection

Small-object detection remains challenging because limited pixels cause information loss and suppress the scale knowledge encoded in pretrained detectors. Existing approaches mainly improve representations through multiscale training, architecture redesign, or parameter adaptation, implicitly assuming that frozen models lack the required capability. We challenge this assumption and hypothesize that small-object knowledge already exists in frozen detectors but remains underactivated and unstable during query evolution. To test this hypothesis, we propose Counterfactual Query-Trajectory Reliability (CQTR), a training-free framework that elicits latent responses through counterfactual scale interventions and interprets candidate reliability from decoder-internal spatial convergence, semantic persistence, and cross-scale conflicts. A small unlabeled training subset selects the appropriate correction mechanism for each model-data stream, without parameter updates or target-domain annotations. Across 27 combinations of nine frozen detectors and three datasets, CQTR consistently improves average precision (AP) and average precision for small objects (APs). Closed-loop analyses further show that scale intervention activates latent responses, trajectory evidence predicts ground-truth support, and unlabeled routing selects the more effective branch. CQTR therefore reframes small-object detection from external scale augmentation to the activation and reliability assessment of latent scale knowledge.

cs.CV

Logos: An Agent Harness on a Cross-Process Bus

Plugin-based agents assemble capabilities at runtime, and the spatiotemporal-composability calculus proves a reversibility guarantee for this assembly. However, the guarantee is carried by a single process, which confines all components, sessions, and recovery records to one failure domain, where a fault spreads past the plugin boundary, and process death interrupts every session the process hosts. Resting only on the hypotheses the calculus already states and the stateless interface of the model call, this paper relaxes the single-process restriction of the calculus to an arbitrary assignment of components and records to processes, gives four sufficient conditions, and proves with Theorem 1, derived from the four lemmas, that the reversibility guarantee holds across processes when these conditions are met. Based on Theorem 1, this paper constructs Logos, a cross-process plugin-based agent in the peer-process and name-routed form of ROS, where a plugin is a process, the router holds only a rebuildable routing table, and the session state needed for recovery lives in an append-only transcript owned by no process. Under one fault on two hundred benchmark tasks across three configurations, the single-process reference lost every session and scored 1.5 percent on the official validator, the MCP configuration kept its sessions while spending 1099 calls on a dead endpoint, and Logos kept every session alive, wasted zero calls, and succeeded on 120 tasks against 102 for both configurations combined. At the mechanism level, eighty sessions terminated at four points of the tool-call cycle all resumed with no repeated action, 3,500 concurrent calls paired with zero violations, and one bus hop cost 1 in 823 of the model's first token. The results show that the reversibility guarantee holds across processes and that assembly itself can leave the host process.

cs.AI

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Suppression), a unified inference-time framework with three complementary components. Cross-Modal Joint Confidence Calibration injects global visual prototypes to calibrate text-driven known-class predictions. Uncertainty-Guided Universal Objectness Enhancement measures classification hesitation from local visual responses to strengthen potential unknown objects. Dynamic Outlier Suppression via Confidence Margin replaces rigid suppression with a margin-aware adjustment that preserves ambiguous out-of-distribution instances. Experiments on the Real-World Detection benchmark demonstrate that, with the OWL-ViT L/14 backbone, CODE achieves 21.7 U-mAP and 40.8 K-mAP in Task 1, surpassing the previous state of the art by 2.6 and 2.3 points, respectively.

cs.CV

Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary Refinement

Vision--language models can identify the correct referent while returning an imprecise bounding box. We study whether a frozen direct-answer model can use its own prediction to allocate one additional localized observation without accessing target annotations at inference. Label-free precision refinement (LFPR) routes predicted-small regions to a higher-resolution pass, re-grounds the expression inside a context crop, admits a candidate only under fixed geometric guards, and returns a fixed coordinate-wise midpoint. We report results across three evidence tiers. On 31,921 retrospective Ref-L4 expressions, LFPR raises mAcc$_{0.5:0.95}$ from 72.947\% to 76.013\% (Acc@0.5 88.531\%$\to$89.725\%, Acc@0.9 55.788\%$\to$61.142\%). A frozen transfer to 30,969 RefCOCO/RefCOCO+/RefCOCOg expressions improves every dataset at Acc@0.5, mAcc, and mean IoU (pooled mAcc $+0.645$, Acc@0.5 $+0.817$), while Acc@0.9 is unchanged overall: routing alone gains $+1.162$ points there, but crop, guards, and fusion give back $-1.192$, offsetting rather than showing no strict-IoU effect. A prospective, image-disjoint Flickr30K Entities evaluation improves every endpoint (mAcc $+0.973$, Acc@0.9 $+1.022$), more strongly under a single-box variant (mAcc $+2.575$, Acc@0.9 $+3.689$). The same operator applied to two released grounding specialists improves every endpoint (Acc@0.9 $+1.569$/$+6.716$ for EGM-4B/8B) at roughly twice the latency, composing with specialist training rather than replacing it. A genuine unguarded control (guard removed from the same candidates) underperforms the incumbent on every metric, showing the guard is load-bearing. Together, these results show that referent selection and boundary precision are partially separable, with different components moving opposing regions of the IoU curve -- behavior a single threshold cannot reveal.

cs.CV

A Systematic Gaia--ZTF Search for Short-Period Blue Compact-Binary Candidates

We present a catalog of 147 short-period (10.34--106.46~min) blue compact-binary candidates, identified by combining Gaia DR3 astrometry and photometry with ZTF DR23 light curves via a Gaia selection, period searches, and machine-learning morphology ranking. Of these, 111 lack prior compact-binary classifications. Multiwavelength data (DESI DR1, GALEX, AllWISE) reveal a heterogeneous sample: on the Gaia colour--magnitude diagram, 52 sources lie on the white-dwarf locus, 69 in the hot-subdwarf region, and 26 are intermediate. Among 26 sources with DESI spectra, only about one third follow the white-dwarf cooling sequence; the rest are more luminous blue stars with white-dwarf-like low-resolution spectra. We highlight a prioritized subset of new white-dwarf-locus candidates for follow-up, including ten with periods below 40~min and none with existing radial-velocity data. Under fiducial binary assumptions, 17 of these newly identified white-dwarf-locus candidates would exceed the adopted LISA signal-to-noise threshold (led by a 37~pc white dwarf), with the count depending on chirp mass (9 for $0.15\,M_\odot$, 17 for $0.3\,M_\odot$, 21 for $0.6\,M_\odot$), assuming orbital modulation. However, for most of the white-dwarf-locus sample, observed modulation amplitudes exceed any plausible ellipsoidal signal by three to five orders of magnitude, implying that rotating magnetic or chemically inhomogeneous single white dwarfs offer a viable alternative that ZTF photometry alone cannot rule out---the catalog includes at least one confirmed case. We release the full 147-source catalog, including periods, Gaia/spectroscopic classifications, harmonic/ellipsoidal diagnostics, and supplementary tables of fiducial GW estimates and UV--IR photometry.

astro-ph.SR

Binary Constraints on the Origin of Nitrogen-rich Field Stars

Recent JWST observations have revealed galaxies with unusually high N/O ratios, suggesting that nitrogen enrichment may be common in intense star-forming environments in the early Universe. In the Milky Way, nitrogen-rich(N-rich) stars in the Galactic field have long served as probes of early Galaxy formation and globular cluster enrichment. However, the identification of binaries among these stars raises the possibility that binary mass transfer could contribute to their origin. In this work, we utilize multi-epoch radial velocities and element abundances from APOGEE DR17 to constrain their formation sites. Among 266 N-rich field stars, 33 exhibit radial velocity variations of $\Delta {\rm RV} > 1\,{\rm km/s}$, including 10 robust spectroscopic binaries identified using the $F_2$ statistic within a well-sampled subset of 46 stars. The resulting close-binary fraction ($21.7\pm6.1\%$) is statistically indistinguishable from that of chemically normal field stars ($18.1\pm0.6\%$), showing no evidence of the excess expected from AGB binary pollution. This is further supported by the absence of correlation between [N/Fe] and [Ce/Fe] and the lack of [C/Fe] enhancement. Crucially, we detect an anti-correlation between binary fraction and [Al/Fe], with strongly Al-enhanced stars ($[\mathrm{Al/Fe}] \gtrsim 0.5$) exhibiting a reduced binary fraction ($< 10\%$). This trend serves as a dynamical fingerprint of high-density environments, consistent with the efficient disruption of binaries via three-body interactions in GC cores. Our results do not support binary mass transfer as the dominant formation channel for N-rich field stars; they are predominantly GC escapees that retain the dynamical memory of their dense birth sites.

astro-ph.GA

Artificial Intelligence for Spatially Reconfigurable Antennas: Movable, Fluid, and Pinching Antenna Systems

Recently, sixth-generation (6G) wireless networks have moved beyond fixed-array designs toward antenna architectures that can adapt their spatial configuration to specific environmental conditions. Movable antenna, fluid antenna, and pinching antenna systems represent this principle in different ways, but they share a common vision: exploiting spatial flexibility as an additional degree of freedom (DoF) to improve communication, sensing, security, and resource efficiency. These new techniques, however, also bring challenging problems, as antenna configuration must be jointly considered with channel acquisition, beamforming, mobility, and network resource management. Therefore, artificial intelligence (AI) has become an important tool for learning fast and adaptive control policies for these highly coupled systems. In this survey, we provide a unified review of AI for spatially reconfigurable antenna systems. We first introduce the basic principles of movable, fluid, and pinching antennas, which is followed by a summary of the latest AI-enabled designs according to their primary optimization objectives. Furthermore, we compare the roles of deep learning (DL), deep reinforcement learning (DRL), multi-agent reinforcement learning (MARL), graph learning, Transformers, large language models (LLMs), and structure-guided learning across different antenna architectures. Finally, we discuss open challenges and future directions toward scalable, robust, and hardware-aware intelligent reconfigurable antenna networks.

eess.SP

Clean-Reference Streaming Detection of Lens Occlusion and Photometric Transitions for Camera Tamper Monitoring

A surveillance camera is an image sensor whose silent physical degradation invalidates every downstream consumer of its data. In-situ integrity alarms for such vision sensors require low false-alarm rates, bounded computation, and diagnosable behavior under nuisance illumination changes. This paper studies a deliberately narrow streaming integrity monitor for two low-cost sensor-fault signatures: texture-collapsing lens occlusion and abrupt photometric scene transition. The detector compares sampled luminance and local-gradient statistics with a clean-only sliding reference, applies coarse-grid structured-light rejection and mode/rapid-brightness suppression, and emits at most one notification per tamper episode. We formalize the decision predicates and derive a consistency rule for when rapid-brightness suppression makes the scene-transition path unreachable. On 320 in-scope controlled sequences, the default state machine attains 0.800 F1 and 0.822 balanced accuracy (significantly better paired correctness than the strongest baseline, though the F1 margin is not statistically resolved); on a magnitude-swept public audit it attains the highest partial AUC under a 5\% false-alarm budget, and a separate extended-stress FPR-constrained sweep reaches 0.925 recall at 0.025 false-positive rate. Public Xiph, Bremen IoT, and UHCTD diagnostics show the fixed predicates preserve low false alarms while recall concentrates inside the declared envelope (UHCTD in-scope covered recall 0.667 versus 0.016 out of scope), and a 9.09-camera-hour verified-negative public audit records zero false alarms. The method is best interpreted as an auditable sensor-health subsystem rather than a universal camera-tamper classifier.

cs.CV

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding

Visual grounding often fails even when the target object is present in the proposal pool, because the language-side referent is bound to the wrong region. We study this binding failure under frozen perception and ask whether an explicit query--region alignment hook, together with a perception-grounded abstention mechanism, can improve faithful grounding without retraining the detector or the vision-language backbone. HKVLM freezes a language-aligned open-vocabulary detector for localization and learns a lightweight hook that maps referential query embeddings to detector proposals in a shared space; a verifier abstains when no region sufficiently supports the query. We prove an exact proposal-level diagnostic decomposition, $(1-\mathrm{SeeErr})(1-\mathrm{SayErr})$, separating proposal-coverage failures from conditional binding failures, and a monotonicity result that characterizes the faithfulness--recall trade-off induced by abstention. Across RefCOCO, RefCOCO+, RefCOCOg, and POPE, HKVLM improves over untrained and trained matched-perception binding controls and substantially reduces hallucination through abstention. Strong coordinate-decoding and end-to-end fine-tuned baselines remain much higher in raw grounding accuracy, and a reasoning-stress set exposes binding as the main current bottleneck. We therefore present HKVLM as a diagnostic and mechanism-level study of query--region binding under frozen perception, not as an absolute localization leader.

cs.CV

Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling

This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detection framework based on biological hippocampal memory modeling. This framework integrates a hippocampal memory network module, HipNet, into the DETR architecture and systematically simulates the anatomical structure and functional organization of hippocampal subregions, including the entorhinal cortex, dentate gyrus, CA3, CA1, and subiculum. Through this design, Hippocampus-DETR realizes pattern separation, pattern completion, importance filtering, and information integration of visual encoding features. During training, different memory submodules are optimized using a layer-wise training strategy, ultimately forming a memory system with memory retrieval and completion capabilities. Experimental results demonstrate that Hippocampus-DETR achieves higher detection accuracy than current mainstream models. More importantly, models equipped with this framework also exhibit excellent generalization ability and data efficiency in tasks such as few-shot image classification, multimodal feature construction, and image restoration. Subsequent experiments further validate the functional necessity and internal interpretability of each memory submodule. This study not only provides a novel object detection framework, but also offers a feasible technical pathway for integrating neurocognitive mechanisms with deep learning models, highlighting its significant value in improving model learning efficiency and task robustness. The project is available at https://github.com/2186cloud/hipnet.

cs.CV

Mass-Orbital Period Distribution of Massive White Dwarfs Formed Through Stable Mass Transfer

White dwarfs (WDs) in binaries can form through either the stable mass-transfer process or common envelope evolution (CEE). Compared to CEE, the stable mass-transfer process can lead to a distinct mass-orbital period ($M_{\mathrm{WD}}-P_{\mathrm{orb}}$) relation. Thus, this relation of WDs contains the information about the evolution channels. We can study the relation in WD binary systems to determine whether their progenitors undergo a CEE. We use the stellar evolution code MESA as our primary computational tool and adopt the quasi-adiabatic criterion to ensure that our models satisfy the conditions for stable mass transfer. Our study considers different mass-transfer schemes, varying metallicities, and the relation for both low-mass and intermediate-mass progenitors. Previous studies have focused on the relation for low-mass progenitors, which cannot explain some long-period, high-mass WD binaries. Our results show that the relations for intermediate-mass progenitors whose cores remain non-degenerate prior to central helium burning can account for the formation channels of long-period and massive WD binaries.

astro-ph.SR

Transformer-Enhanced Reinforcement Learning: Fundamentals and Applications in Communication Networks

Reinforcement Learning (RL) has long been a powerful solution to various problems in communication networks. However, traditional RL models still face with several limitations. Not only do they rely on large numbers of interactions with the environment, but they are also limited in terms of modeling long-term relationships and tackling partial observability. In recent years, the Transformer model has demonstrated the ability to enhance RL models, allowing them to overcome these issues. Particularly, the self-attention mechanism within the Transformer enables efficient modeling of long-range dependencies and global correlations, as well as accelerates training processes and handles heterogeneous data modalities. In this paper, we present a comprehensive survey of Transformer-based RL algorithms and their applications in communication networks. Specifically, the paper provides the mathematical background of RL and Transformer architectures, along with insights into key issues such as resource allocation, computation offloading, routing, and trajectory control, and network security. We conclude the paper by discussing challenges, open issues, and notable future research directions, including Transformer-enhanced DRL algorithms for semantic communication and network optimization.

eess.SP

From Denoising to Decision Making: A Survey on Diffusion Model-Enabled Deep Reinforcement Learning for Wireless Networks

Deep reinforcement learning (DRL) has long been a promising solution for sequential resource management in wireless networks. However, conventional DRL methods are fundamentally limited by their reliance on unimodal policy distributions, inefficient exploration in high-dimensional action spaces, and poor adaptability to dynamic and heterogeneous environments. Meanwhile, diffusion models (DMs) as one of the most powerful families of generative AI have demonstrted remarkable capabilities in modeling complex, multi-modal data distributions across diverse domains. The integration of DMs and DRL has opened a new and rapidly growing research direction, in which DM-enabled policies substantially enhance decision quality by capturing the complex, discontinuous, and multimodal action structures inherent in wireless resource management. In this paper, we present a comprehensive survey of DM-enabled DRL algorithms and their applications for various issues in wireless networks. Particularly, we first provide the theoretical background of DM and present different DM-enabled DRL algorithms. We then systematically review applications of DM-enabled DRL for across computation offloading in mobile edge computing, UAV-assisted, vehicular, and AIGC-driven systems, as well as wireless resource allocation, physical-layer security, and robotics and UAV planning. We conclude the paper by higlight future research directions.

eess.SP

Design, Testing, and Commissioning of the Sun Yat-sen University (SYSU) 80 cm Infrared Telescope

The Sun Yat-sen University (SYSU) 80 cm telescope is a new generation near-infrared (NIR) facility in China dedicated to time-domain astronomy, while also serving as a testbed for emerging NIR cameras. Commissioned in October 2024 at the 4100 m Lenghu site on the Tibetan Plateau in China, the telescope adopts a reflective Cassegrain design with two Nasmyth foci for J and K bands. The J band imaging system, initially equipped with a 640 x 512 off-the-shelf InGaAs camera (INS Mars640) and upgraded in June 2025 to a 1280 x 1024 science-grade, deeply cooled camera (YNAOIR), achieves background-limited performance with a dark current of ~ 14 e-/s/pix and a readout noise of ~ 11 e-. The system reaches a limiting magnitude of J ~ 17 mag (Vega system) in single 20 s exposures and depths of J ~ 19.4 mag with stacked 30 minute exposures. For a variable with J ~ 14 mag during on-sky tests, the system delivers millimagnitude-level photometric precision. Since commissioning, the telescope observed transients such as gamma-ray bursts (GRBs), supernovae and comets, variables including active galactic nuclei (AGNs), high-redshift quasars (z > 6), and brown dwarfs, as well as deep-field imaging reaching J ~ 20.5 mag. This validates the feasibility of using InGaAs cameras for astronomical observations, encouraging other institutions to develop dedicated infrared telescopes or integrate infrared cameras into existing optical telescopes.

astro-ph.IM

Fates of the sub-stellar objects (FOSSO) II. Evidence for Suppression of Metal Pollution in White Dwarfs by Close Substellar Companions

Approximately 25--50\% of white dwarfs (WDs) exhibit metal absorption lines in their photospheres, interpreted as evidence of ongoing/recent accretion of planetary debris from remnant systems. Previous theoretical studies have suggested that massive, close-in substellar companion may prevent delivery of larger bodies via dynamical interactions, thereby reducing white-dwarf pollution. However, no conclusive observational evidence has yet been established to confirm such a protective effect. In this work, based on a sample of 17 white dwarf-substellar companion (1--75 $M_{\rm J}$) systems with reliable spectroscopic classifications, we find that white dwarfs hosting close substellar companions (orbital period $P < 5$ d) exhibit a metal-pollution fraction of $7.7^{+11.3}_{-4.0}\%$, which is suppressed by a factor of $5.75^{+3.24}_{-1.94}$ (corresponding to a protection efficiency of $87.2^{+3.4}_{-9.2}\%$) relative to single white dwarfs with a confidence level of 99.96\%. In contrast, white dwarfs with wider companions show a metal-pollution fraction of approximately $25.0^{+24.0}_{-12.8}\%$, comparable to that of single white dwarf systems. To interpret these results, we perform ensembles of N-body integrations and demonstrate that massive close-in substellar companions are capable of clearing 80\%--90\% of small-body contaminants. The good consistency between the observational statistics and dynamical simulations provides strong evidence for suppressed metal pollution in white dwarfs with close companions, and offers insights into the long-term dynamical evolution of WD remanent systems.

astro-ph.SR

Sub-Neptunes Show a Stronger Correlation with Cold Jupiters than Super-Earths Especially in Metal-rich Systems

Correlations between the inner small planets and cold giants encodes the formation and evolution of planetary systems. It remains unclear if the correlation differs on the two sides of the radius valley. In this work, we compute the conditional frequency of cold Jupiters in systems with only inner sub-Neptunes $P(\rm CJ|SN)$ and those with only inner super-Earths $P(\rm CJ|SE)$. We find that, around transiting sample around metal-rich stars, $P(\rm CJ|SN, [Fe/H]>0)$ and $P(\rm CJ|SE, [Fe/H]>0)$ are $42.6^{+10.6}_{-9.9}\%$ and $14.5^{+12.7}_{-6.9}\%$. Comparing with the field giant frequency ($14.3^{+2.0}_{-1.8}\%$), we show that inner sub-Neptunes and cold Jupiters exhibit a significant positive correlation for metal-rich systems with a confidence level of 99.95\%, whereas this correlation is absent for systems with super-Earths. We also consider a homogeneous Kepler-Keck subsample and derive similar results, with $P(\rm CJ|SN, [Fe/H]>0)$ of $45.8^{+18.6}_{-16.3}\%$ and $P(\rm CJ|SE, [Fe/H]>0)$ of $13.3^{+17.0}_{-6.8}\%$. Radial velocity sample shows consistent results, with metal-rich systems hosting massive inner planets exhibiting a strong positive correlation (confidence level of 99.11\%) with outer cold Jupiters ($P(\rm CJ|M_{p}>10M_\oplus, [Fe/H]>0) = 34.6^{+11.0}_{-9.1}\%$). These results can be naturally understood since metal-rich disks are expected to more efficiently produce both outer cold Jupiters and inner planets with larger radii and masses. Our findings highlight the critical role of stellar metallicity in shaping planetary architectures, particularly for large/massive planets.

astro-ph.EP

PPEDCRF: Dynamic-CRF-Guided Selective Perturbation for Background-Based Location Privacy in Video Sequences

We propose PPEDCRF, a calibrated selective perturbation framework that protects \emph{background-based location privacy} in released video frames against gallery-based retrieval attackers. Even after GPS metadata are stripped, an adversary can geolocate a frame by matching its background visual cues to geo-tagged reference imagery; PPEDCRF mitigates this threat by estimating location-sensitive background regions with a dynamic conditional random field (DCRF), rescaling perturbation strength with a normalized control penalty (NCP), and injecting Gaussian noise only inside the inferred regions via a DP-style calibration rule. On a controlled paired-scene retrieval benchmark with eight attacker backbones and three noise seeds, PPEDCRF reduces ResNet18 Top-1 retrieval accuracy from 0.667 to $0.361\pm0.127$ at $\sigma_0=8$ while preserving $36.14\,$dB PSNR -- an ${\approx}6\,$dB quality advantage over global Gaussian noise. Transfer across the eight-backbone seed-averaged benchmark is broadly supportive (23 of 24 backbone-gallery cells show negative $\Delta$), while appendix-scale confirmation identifies MixVPR as a remaining adverse-transfer exception. Matched-operating-point analysis shows that PPEDCRF and global Gaussian noise converge in Top-1 privacy at equal utility, so the practical benefit is spatially concentrated perturbation that preserves higher visual quality at any given noise scale rather than stronger matched-utility privacy. Code: https://github.com/mabo1215/PPEDCRF

cs.CV

From Exploration to Specification: LLM-Based Property Generation for Mobile App Testing

Mobile apps often suffer from functional bugs that do not cause crashes but instead manifest as incorrect behaviors under specific user interactions. Such bugs are difficult to detect automatically because they often lack explicit test oracles. Property-based testing can effectively expose them by checking intended behavioral properties under diverse interactions. However, its use largely depends on manually written properties, whose construction is difficult and expensive, limiting its practical use for mobile apps. To address this limitation, we propose PropGen, an automated approach for generating properties for Android apps. However, this task is challenging for two reasons: app functionalities are often hard to systematically uncover and execute, and properties are difficult to derive accurately from observed behaviors. To this end, PropGen performs functionality-guided exploration to collect behavioral evidence from app executions, synthesizes properties from the collected evidence, and refines imprecise properties based on testing feedback. We implemented PropGen and evaluated it on 12 real-world Android apps. The results show that PropGen can effectively identify and execute valid app functionalities, generate valid properties, and repair most imprecise ones. Across all apps, PropGen identified 1,210 valid functionalities and correctly executed 977 of them, compared with 491 and 187 for the baseline. It generated 985 properties, 912 of which were valid, and repaired 118 of 127 imprecise ones exposed during testing. With the resulting properties, we found 25 previously unknown functional bugs in the latest versions of the subject apps, many of which were missed by existing functional testing techniques.

cs.SE