SearcharxivSearch

arXiv subjects

Jingxiao Ma

Publications and source records attributed to Jingxiao Ma.

10 recordsLinked to original sources

SetMIR: Multi-Interest Retrieval as Set Prediction

Embedding-based retrieval is at the core of industrial recommender systems, but a single user embedding is often too limited to capture a user's diverse interests. Multi-interest retrieval addresses this by using multiple user embeddings, yet existing methods still suffer from two issues: interest collapse, where different embeddings learn the same interest, and static dispatch, where serving uses a fixed retrieval budget even when some embeddings are unnecessary. We propose SetMIR, which treats multi-interest retrieval as a set prediction problem. SetMIR encodes a user's behavior history with a transformer and uses K learnable queries to decode a set of user interests, each producing a retrieval embedding and a presence score. During training, Hungarian matching assigns targets to queries one-to-one, so matched queries learn distinct interests and the presence head learns which queries are active. At serving time, SetMIR uses presence scores and query-level Non-Maximum Suppression (NMS) to issue only active, non-redundant ANN queries. On Snap's Dynamic Product Ads (DPA) data, SetMIR outperforms four learned multi-interest retrievers on every metric while issuing 33% fewer ANN queries per request. Deployed as a new retrieval source in the DPA production stack, SetMIR lifts overall CVR by 3.1%, while lifting CTR by 44% and CVR by 51% over the item-to-item retrieval source with the same item embeddings, ANN index, and retrieval quota.

cs.IR

CAMIE: Co-Engagement-Aware Multimodal Item Embeddings for Snap Dynamic Product Ads Retrieval

Item-to-item (I2I) retrieval is a core primitive in large-scale recommendation and advertising systems. In production Snap Dynamic Product Ads (DPA), I2I retrieval faces two challenges: separate visual, textual, and multimodal encoders fragment the retrieval stack, and content-only training does not align embeddings with the co-engagement behavior that drives downstream conversions. We present CAMIE, a co-engagement-aware multimodal item embedding framework for Snap DPA retrieval. CAMIE builds on LLM/MLLM backbones, using their native multimodal interfaces to represent item images and metadata in a shared embedding space. It then fine-tunes the backbone on co-engaged item pairs mined from user journeys with a symmetric in-batch InfoNCE objective. Offline, CAMIE outperforms the strongest commercial multimodal embedding model on Recall@10 and serves text-only retrieval from the same checkpoint with minimal quality loss. Online, CAMIE serves as a drop-in replacement for two deployed content-based I2I encoders, delivering +0.390% CTR / +10.832% CVR over the multimodal control, +18.958% CTR / +13.12% CVR over the text control, and +0.211% CTR / +1.911% CVR on overall DPA traffic. CAMIE is deployed in production.

cs.IR

EGR: Embedding-Native Generative Retrieval with a Shared LLM

Generative retrieval is increasingly popular in large-scale recommendation and advertising systems, yet current methods introduce practical complications. Semantic-ID methods rely on quantization, mutable identifier vocabularies, and token-to-item grounding; embedding-based pipelines train the item encoder separately from the query generator, which limits user-item alignment. We propose EGR, an Embedding-Native Generative Retrieval framework for recommendation and advertising. EGR uses a single shared LLM to learn item representations from item metadata and user representations from interaction histories in one embedding space. Items are indexed directly as dense vectors, and user histories are encoded as dense retrieval queries. Joint contrastive training groups related items and aligns queries with their target items. We evaluate EGR on public benchmarks, industrial data, and live deployment. EGR outperforms published baselines on Amazon Reviews; on Snap DPA, it scales with data, handles cold-start items, and benefits from multimodal input. In production, EGR delivers a +2.91% conversion-rate lift, simplifying system design while improving retrieval quality and ad performance.

cs.IR

SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads

Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories). While Large Language Models (LLMs) capture semantic intent better than traditional embedding models, deploying them at scale introduces prohibitive inference costs and lexical mismatch issues. Through controlled experiments on millions of users, we demonstrate a critical retrieval decomposition: rule-generated queries excel at retargeting on a lexical BM25 index, while LLM-generated queries excel at prospecting on a dense ANN index. Building on this, we propose SMART (SeMantic-aware Adaptive ReTrieval). To manage costs, a lightweight quality gate identifies coverage gaps in initial keyword results, adaptively routing only the ~10% of users who benefit from semantic prospecting to the LLM path. Offline evaluation demonstrates that this gated approach captures the bulk of semantic prospecting gains in Relevance Score while maintaining competitive re-targeting performance at a 90% reduction in LLM costs. Finally, in a 2-week online A/B test at Snap, SMART improved the ad conversion rate by +27.6% over a strong embedding-based baseline.

cs.IR

FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision

Backpropagation has been the cornerstone of neural network training for decades, yet its inefficiencies in time and energy consumption limit its suitability for resource-constrained edge devices. While low-precision neural network quantization has been extensively researched to speed up model inference, its application in training has been less explored. Recently, the Forward-Forward (FF) algorithm has emerged as a promising alternative to backpropagation, replacing the backward pass with an additional forward pass. By avoiding the need to store intermediate activations for backpropagation, FF can reduce memory footprint, making it well-suited for embedded devices. This paper presents an INT8 quantized training approach that leverages FF's layer-by-layer strategy to stabilize gradient quantization. Furthermore, we propose a novel "look-ahead" scheme to address limitations of FF and improve model accuracy. Experiments conducted on NVIDIA Jetson Orin Nano board demonstrate 4.6% faster training, 8.3% energy savings, and 27.0% reduction in memory usage, while maintaining competitive accuracy compared to the state-of-the-art.

cs.LG

Approximate Logic Synthesis Using BLASYS

Approximate computing is an emerging paradigm where design accuracy can be traded for improvements in design metrics such as design area and power consumption. In this work, we overview our open-source tool, BLASYS, for synthesis of approximate circuits using Boolean Matrix Factorization (BMF). In our methodology the truth table of a given circuit is approximated using BMF to a controllable approximation degree, and the results of the factorization are used to synthesize the approximate circuit output. BLASYS scales up the computations to large circuits through the use of partition techniques, where an input circuit is partitioned into a number of interconnected subcircuits and then a design-space exploration technique identifies the best order for subcircuit approximations. BLASYS leads to a graceful trade-off between accuracy and full circuit complexity as measured by design area. Using an open-source design flow, we extensively evaluate our methodology on a number of benchmarks, where we demonstrate that the proposed methodology can achieve on average 48.14% in area savings, while introducing an average relative error of 5%.

cs.AR

MetRex: A Benchmark for Verilog Code Metric Reasoning Using LLMs

Large Language Models (LLMs) have been applied to various hardware design tasks, including Verilog code generation, EDA tool scripting, and RTL bug fixing. Despite this extensive exploration, LLMs are yet to be used for the task of post-synthesis metric reasoning and estimation of HDL designs. In this paper, we assess the ability of LLMs to reason about post-synthesis metrics of Verilog designs. We introduce MetRex, a large-scale dataset comprising 25,868 Verilog HDL designs and their corresponding post-synthesis metrics, namely area, delay, and static power. MetRex incorporates a Chain of Thought (CoT) template to enhance LLMs' reasoning about these metrics. Extensive experiments show that Supervised Fine-Tuning (SFT) boosts the LLM's reasoning capabilities on average by 37.0\%, 25.3\%, and 25.7\% on the area, delay, and static power, respectively. While SFT improves performance on our benchmark, it remains far from achieving optimal results, especially on complex problems. Comparing to state-of-the-art regression models, our approach delivers accurate post-synthesis predictions for 17.4\% more designs (within a 5\% error margin), in addition to offering a 1.7x speedup by eliminating the need for pre-processing. This work lays the groundwork for advancing LLM-based Verilog code metric reasoning.

cs.AR

RUCA: RUntime Configurable Approximate Circuits with Self-Correcting Capability

Approximate computing is an emerging computing paradigm that offers improved power consumption by relaxing the requirement for full accuracy. Since real-world applications may have different requirements for design accuracy, one trend of approximate computing is to design runtime quality-configurable circuits, which are able to operate under different accuracy modes with different power consumption. In this paper, we present a novel framework RUCA which aims to approximate an arbitrary input circuit in a runtime configurable fashion. By factorizing and decomposing the truth table, our approach aims to approximate and separate the input circuit into multiple configuration blocks which support different accuracy levels, including a corrector circuit to restore full accuracy. By activating different blocks, the approximate circuit is able to operate at different accuracy-power configurations. To improve the scalability of our algorithm, we also provide a design space exploration scheme with circuit partitioning to navigate the search space of possible approximations of subcircuits during design time. We thoroughly evaluate our methodology on a set of benchmarks and compare against another quality-configurable approach, showcasing the benefits and flexibility of RUCA. For 3-level designs, RUCA saves power consumption by 36.57% within 1% error and by 51.32% within 2% error on average.

cs.AR

Robust Secure Transmission Design for IRS-Assisted mmWave Cognitive Radio Networks

Cognitive radio networks (CRNs) and millimeter wave (mmWave) communications are two major technologies to enhance the spectrum efficiency (SE). Considering that the SE improvement in the CRNs is limited due to the interference temperature imposed on the primary user (PU), and the severe path loss and high directivity in mmWave communications make it vulnerable to blockage events, we introduce an intelligent reflecting surface (IRS) into mmWave CRNs. Due to the estimation mismatch and the passivity of Eavesdroppers (Eves), perfect channel state information (CSI) of wiretap links is challenging to obtain, which promotes our research on robust secure beamforming (BF) design in the IRS-assisted mmWave CRNs. This paper considers the collaborate scenario of Eves, which allows us to investigate the BF design in the harsh eavesdropping environment. Specifically, by using a uniform linear array (ULA) at the cognitive base station (CBS) and a uniform planar array (UPA) at the IRS, and supposing that imperfect CSIs of angle-of-departures for wiretap links are known, we formulate a constrained problem to maximize the worst-case achievable secrecy rate (ASR) of the secondary user (SU) by jointly designing the transmit BF at the CBS and reflect BF at the IRS. To solve the non-convex problem with coupled variables, an efficient alternating optimization algorithm is proposed. Finally, simulation results indicate that the ASR performance of our proposed algorithm has a small gap with that of the optimal solution with perfect CSI compared with the other benchmarks.

cs.IT

Secure and Energy Efficient Transmission for IRS-Assisted Cognitive Radio Networks

The spectrum efficiency (SE) and security of the secondary users (SUs) in the cognitive radio networks (CRNs) have become two main issues due to the limitation interference to the primary users (PUs) and the shared spectrum with the PUs. Intelligent reflecting surface (IRS) has been recently proposed as a revolutionary technique which can help to enhance the SE and physical layer security of wireless communications. This paper investigates the application of IRS in an underlay CRN, where a multi-antenna cognitive base station (CBS) utilizes spectrum assigned to the PU to communicate with a SU via IRS in the presence of multiple coordinated eavesdroppers (Eves). To achieve the trade-off between the secrecy rate (SR) and energy consumption, we investigate the secrecy energy efficiency (SEE) maximization problem by jointly designing the transmit beamforming at the CBS and the reflect beamforming at the IRS. To solve the non-convex problem with coupled variables, we propose an iterative alternating optimization algorithm to solve the sub-problems alternately, by utilizing an iterative penalty function based algorithm for sub-problem 1 and the difference of two-convex functions method for sub-problem 2. Furthermore, we provide a second-order-cone-programming (SOCP) approximation approach to reduce the computational complexity. Finally, the simulation results demonstrate that IRS can help significantly improve the SE and enhance the physical layer security in the CRNs. Moreover, the effectiveness and superiority of our proposed algorithm in achieving the trade-off between the SR and energy consumption are verified.

cs.IT