SearcharxivSearch

arXiv subjects

Mohammad Saleh

Publications and source records attributed to Mohammad Saleh.

At least 19 recordsLinked to original sources

Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition

Reliable fall detection in elderly care requires monitoring systems that are not only accurate but also capable of producing stable, interpretable explanations of motion dynamics, a requirement that existing post hoc explainability methods rarely satisfy when applied to sequential biosignals. This study introduces a lightweight framework for skeleton-based fall detection that combines a Long Short-Term Memory (LSTM) model with a temporally stabilized attribution mechanism. We propose Temporal SHAP (T-SHAP), which treats frame-wise SHAP attributions as a temporal signal and applies a linear smoothing operator to reduce high-frequency variance. From a signal processing perspective, this operation is analogous to low-pass filtering, enabling the extraction of consistent temporal patterns while preserving the theoretical properties of Shapley-based attributions. Experiments conducted on the NTU RGB+D dataset demonstrate that the proposed approach achieves 94.3% classification accuracy with an end-to-end latency below 25 ms, supporting real-time applicability. Quantitative evaluation using perturbation-based faithfulness metrics shows that T-SHAP improves attribution reliability compared to standard SHAP (AUP: 0.91 vs. 0.89) and Grad-CAM (0.82), while also reducing temporal variance in the attribution signals. The resulting explanations highlight biomechanically relevant motion patterns, such as lower-limb instability and changes in trunk posture, which are consistent with known characteristics of fall events. The resulting framework is computationally lightweight, requires no additional model training, and produces explanations that are both temporally stable and biomechanically meaningful, properties directly relevant to the reliability demands of AI-assisted clinical monitoring.

cs.CV

Uniformly S-pseudo-projective modules

In this paper, we introduce the notion of uniformly S-pseudo-projective (u-S-pseudo-projective) modules as a generalization of u-S-projective modules. Let R be a ring and S a multiplicative subset of R. An R-module P is said to be u-S-pseudo-projective if for any submodule K of P, there is s\in S such that for any u-S-epimorphism f:P\to \frac{P}{K}, sf can be lifted to an endomorphism g:P\to P. We prove that an R-module M is u-S-quasi-projective if and only if M\oplus M is u-S-pseudo-projective. Also, we prove that if A\oplus B is u-S-pseudo-projective, then any u-S-epimorphism f:A\to B u-S-splits. We give characterizations of certain classes of rings, such as u-S-semisimple and strongly S-perfect rings.

math.AC

On 1-absorbing prime and weakly 1-absorbing prime subsemimodules

In this paper, we introduce the concepts of 1-absorbing prime and weakly 1-absorbing prime subsemimodules over commutative semirings. Let S be a commutative semiring with 1 \neq 0 and M an S-semimodule. A proper subsemimodule N of M is called 1-absorbing prime (weakly 1-absorbing prime) if, for all nonunits a, b \in S and m \in M, abm \in N (0 \neq abm \in N) implies ab \in (N :_{S} M) or m \in N. We study many properties of these concepts. For example, we show that a proper subsemimodule N of M is 1-absorbing prime if and only if for all proper ideals I, J of S and subsemimodule K of M with IJK \subseteq N, either IJ \subseteq (N:_{S} M) or K \subseteq N. Also, we prove that a proper subtractive subsemimodule N of M is weakly 1-absorbing prime if and only if for all proper ideals I, J of S and subsemimodule K of M with 0 \neq IJK \subseteq N, either IJ \subseteq (N:_{S} M) or K \subseteq N.

math.AC

Uniformly S-pseudo-injective modules

This paper introduces the notion of uniformly-S-pseudo-injective (u-S-pseudo-injective) modules as a generalization of u-S-injective modules. Let R be a ring and S a multiplicative subset of R. An R-module E is said to be u-S-pseudo-injective if for any submodule K of E, there is s in S such that for any u-S-monomorphism f : K \to E, sf can be extended to an endomorphism g : E \to E. Several properties of this notion are studied. For example, we show that an R-module M is u-S-quasi-injective if and only if M \oplus M is u-S-pseudo-injective. Two classes of rings related to the class of QI-rings are introduced and characterized.

math.AC

Uniformly S-essential submodules and uniformly S-injective uniformly S-envelopes

In this paper, we introduce the notion of uniformly S-essential (u-S-essential) submodules. Let R be a commutative ring, S a multiplicative subset of R, and M an R-module. A submodule N of M is said to be u-S-essential in M if for any submodule L of M, N \cap L is u-S-torsion implies L is u-S-torsion. Several properties of this notion are studied. We also introduce the notions of u-S-uniform modules and u-S-injective u-S-envelopes and characterize them in terms of u-S-essential submodules.

math.AC

Uniformly S-projective relative to a module and its dual

In this article, we introduce the notion of uniformly S-projective (u-S-projective) relative to a module. Let S be a multiplicative subset of a ring R and M an R-module. An R-module P is said to be u-S-projective relative to M if for any u-S-epimorphism f : M \to N, the induced map HomR(P, f ): HomR(P, M ) \to HomR(P, N ) is a u-S-epimorphism. Dually, we also introduce u-S-injective relative to a module. Some properties of these notions are discussed. Several characterizations of u-S-semisimple modules are given in terms of these notions. The notions of u-S-quasi-projective and u-S-quasi-injective modules are also introduced, and some of their properties are discussed.

math.AC

Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks

Objective: This review explores the trustworthiness of multimodal artificial intelligence (AI) systems, specifically focusing on vision-language tasks. It addresses critical challenges related to fairness, transparency, and ethical implications in these systems, providing a comparative analysis of key tasks such as Visual Question Answering (VQA), image captioning, and visual dialogue. Background: Multimodal models, particularly vision-language models, enhance artificial intelligence (AI) capabilities by integrating visual and textual data, mimicking human learning processes. Despite significant advancements, the trustworthiness of these models remains a crucial concern, particularly as AI systems increasingly confront issues regarding fairness, transparency, and ethics. Methods: This review examines research conducted from 2017 to 2024 focusing on forenamed core vision-language tasks. It employs a comparative approach to analyze these tasks through the lens of trustworthiness, underlining fairness, explainability, and ethics. This study synthesizes findings from recent literature to identify trends, challenges, and state-of-the-art solutions. Results: Several key findings were highlighted. Transparency: Explainability of vision language tasks is important for user trust. Techniques, such as attention maps and gradient-based methods, have successfully addressed this issue. Fairness: Bias mitigation in VQA and visual dialogue systems is essential for ensuring unbiased outcomes across diverse demographic groups. Ethical Implications: Addressing biases in multilingual models and ensuring ethical data handling is critical for the responsible deployment of vision-language systems. Conclusion: This study underscores the importance of integrating fairness, transparency, and ethical considerations in developing vision-language models within a unified framework.

cs.CR

Building Math Agents with Multi-Turn Iterative Preference Learning

Recent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning. While current methods focus on synthetic data generation and Supervised Fine-Tuning (SFT), this paper studies the complementary direct preference learning approach to further improve model performance. However, existing direct preference learning algorithms are originally designed for the single-turn chat task, and do not fully address the complexities of multi-turn reasoning and external tool integration required for tool-integrated mathematical reasoning tasks. To fill in this gap, we introduce a multi-turn direct preference learning framework, tailored for this context, that leverages feedback from code interpreters and optimizes trajectory-level preferences. This framework includes multi-turn DPO and multi-turn KTO as specific implementations. The effectiveness of our framework is validated through training of various language models using an augmented prompt set from the GSM8K and MATH datasets. Our results demonstrate substantial improvements: a supervised fine-tuned Gemma-1.1-it-7B model's performance increased from 77.5% to 83.9% on GSM8K and from 46.1% to 51.2% on MATH. Similarly, a Gemma-2-it-9B model improved from 84.1% to 86.3% on GSM8K and from 51.0% to 54.5% on MATH.

cs.LG

RRM: Robust Reward Model Training Mitigates Reward Hacking

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on RewardBench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%.

cs.CL

On Cellular Automata

Cellular automata are a fundamental computational model with applications in mathematics, computer science, and physics. In this work, we explore the study of cellular automata to cases where the universe is a group, introducing the concept of \( ϕ\)-cellular automata. We establish new theoretical results, including a generalized Uniform Curtis-Hedlund Theorem and linear \( ϕ\)-cellular automata. Additionally, we define the covering map for \( ϕ\)-cellular automata and investigate its properties. Specifically, we derive results for quotient covers when the universe of the automaton is a circulant graph. This work contributes to the algebraic and topological understanding of cellular automata, paving the way for future exploration of different types of covers and their applications to broader classes of graphs and dynamical systems.

math.GR

LiPO: Listwise Preference Optimization through Learning-to-Rank

Aligning language models (LMs) with curated human feedback is critical to control their behaviors in real-world applications. Several recent policy optimization methods, such as DPO and SLiC, serve as promising alternatives to the traditional Reinforcement Learning from Human Feedback (RLHF) approach. In practice, human feedback often comes in a format of a ranked list over multiple responses to amortize the cost of reading prompt. Multiple responses can also be ranked by reward models or AI feedback. There lacks such a thorough study on directly fitting upon a list of responses. In this work, we formulate the LM alignment as a \textit{listwise} ranking problem and describe the LiPO framework, where the policy can potentially learn more effectively from a ranked list of plausible responses given the prompt. This view draws an explicit connection to Learning-to-Rank (LTR), where most existing preference optimization work can be mapped to existing ranking objectives. Following this connection, we provide an examination of ranking objectives that are not well studied for LM alignment with DPO and SLiC as special cases when list size is two. In particular, we highlight a specific method, LiPO-$λ$, which leverages a state-of-the-art \textit{listwise} ranking objective and weights each preference pair in a more advanced manner. We show that LiPO-$λ$ can outperform DPO variants and SLiC by a clear margin on several preference alignment tasks with both curated and real rankwise preference data.

cs.CL

Further Notes on Tightness

Tight and essentially tight modules generalize weakly injective modules. Essential tightness requires embeddings to be essential. This restriction makes the two notions totally different. In this note, we investigate cases when those two notions are the same. Moreover, we look at the cases when essentiallity is imposed only on one of the embeddings rather than both. This allows defining a special class of tight and essentially tight modules and a generalization of both.

math.RA

Statistical Rejection Sampling Improves Preference Optimization

Improving the alignment of language models with human preferences remains an active research challenge. Previous approaches have primarily utilized Reinforcement Learning from Human Feedback (RLHF) via online RL methods such as Proximal Policy Optimization (PPO). Recently, offline methods such as Sequence Likelihood Calibration (SLiC) and Direct Preference Optimization (DPO) have emerged as attractive alternatives, offering improvements in stability and scalability while maintaining competitive performance. SLiC refines its loss function using sequence pairs sampled from a supervised fine-tuned (SFT) policy, while DPO directly optimizes language models based on preference data, foregoing the need for a separate reward model. However, the maximum likelihood estimator (MLE) of the target optimal policy requires labeled preference pairs sampled from that policy. DPO's lack of a reward model constrains its ability to sample preference pairs from the optimal policy, and SLiC is restricted to sampling preference pairs only from the SFT policy. To address these limitations, we introduce a novel approach called Statistical Rejection Sampling Optimization (RSO) that aims to source preference data from the target optimal policy using rejection sampling, enabling a more accurate estimation of the optimal policy. We also propose a unified framework that enhances the loss functions used in both SLiC and DPO from a preference modeling standpoint. Through extensive experiments across three diverse tasks, we demonstrate that RSO consistently outperforms both SLiC and DPO on evaluations from both Large Language Model (LLM) and human raters.

cs.CL

Improving the Robustness of Summarization Models by Detecting and Removing Input Noise

The evaluation of abstractive summarization models typically uses test data that is identically distributed as training data. In real-world practice, documents to be summarized may contain input noise caused by text extraction artifacts or data pipeline bugs. The robustness of model performance under distribution shift caused by such noise is relatively under-studied. We present a large empirical study quantifying the sometimes severe loss in performance (up to 12 ROUGE-1 points) from different types of input noise for a range of datasets and model sizes. We then propose a light-weight method for detecting and removing such noise in the input during model inference without requiring any extra training, auxiliary models, or even prior knowledge of the type of noise. Our proposed approach effectively mitigates the loss in performance, recovering a large fraction of the performance drop, sometimes as large as 11 ROUGE-1 points.

cs.CL

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Learning from human feedback has been shown to be effective at aligning language models with human preferences. Past work has often relied on Reinforcement Learning from Human Feedback (RLHF), which optimizes the language model using reward scores assigned from a reward model trained on human preference data. In this work we show how the recently introduced Sequence Likelihood Calibration (SLiC), can also be used to effectively learn from human preferences (SLiC-HF). Furthermore, we demonstrate this can be done with human feedback data collected for a different model, similar to off-policy, offline RL data. Automatic and human evaluation experiments on the TL;DR summarization task show that SLiC-HF significantly improves supervised fine-tuning baselines. Furthermore, SLiC-HF presents a competitive alternative to the PPO RLHF implementation used in past work while being much simpler to implement, easier to tune and more computationally efficient in practice.

cs.CL

Out-of-Distribution Detection and Selective Generation for Conditional Language Models

Machine learning algorithms typically assume independent and identically distributed samples in training and at test time. Much work has shown that high-performing ML classifiers can degrade significantly and provide overly-confident, wrong classification predictions, particularly for out-of-distribution (OOD) inputs. Conditional language models (CLMs) are predominantly trained to classify the next token in an output sequence, and may suffer even worse degradation on OOD inputs as the prediction is done auto-regressively over many steps. Furthermore, the space of potential low-quality outputs is larger as arbitrary text can be generated and it is important to know when to trust the generated output. We present a highly accurate and lightweight OOD detection method for CLMs, and demonstrate its effectiveness on abstractive summarization and translation. We also show how our method can be used under the common and realistic setting of distribution shift for selective generation (analogous to selective prediction for classification) of high-quality outputs, while automatically abstaining from low-quality ones, enabling safer deployment of generative language models.

cs.CL

Calibrating Sequence likelihood Improves Conditional Language Generation

Conditional language models are predominantly trained with maximum likelihood estimation (MLE), giving probability mass to sparsely observed target sequences. While MLE trained models assign high probability to plausible sequences given the context, the model probabilities often do not accurately rank-order generated sequences by quality. This has been empirically observed in beam search decoding as output quality degrading with large beam sizes, and decoding strategies benefiting from heuristics such as length normalization and repetition-blocking. In this work, we introduce sequence likelihood calibration (SLiC) where the likelihood of model generated sequences are calibrated to better align with reference sequences in the model's latent space. With SLiC, decoding heuristics become unnecessary and decoding candidates' quality significantly improves regardless of the decoding method. Furthermore, SLiC shows no sign of diminishing returns with model scale, and presents alternative ways to improve quality with limited training and inference budgets. With SLiC, we exceed or match SOTA results on a wide range of generation tasks spanning abstractive summarization, question generation, abstractive question answering and data-to-text generation, even with modest-sized models.

cs.CL

Assessing The Factual Accuracy of Generated Text

We propose a model-based metric to estimate the factual accuracy of generated text that is complementary to typical scoring schemes like ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and BLEU (Bilingual Evaluation Understudy). We introduce and release a new large-scale dataset based on Wikipedia and Wikidata to train relation classifiers and end-to-end fact extraction models. The end-to-end models are shown to be able to extract complete sets of facts from datasets with full pages of text. We then analyse multiple models that estimate factual accuracy on a Wikipedia text summarization task, and show their efficacy compared to ROUGE and other model-free variants by conducting a human evaluation study.

cs.CL