SearcharxivSearch

arXiv subjects

Yongquan Hu

Publications and source records attributed to Yongquan Hu.

At least 19 recordsLinked to original sources

On the constituents of the mod $p$ cohomology of Shimura curves

Let $p$ be a prime number and $K$ a finite unramified extension of $\mathbb{Q}_p$. When $p$ is large enough with respect to $[K:\mathbb{Q}_p]$ and under mild genericity assumptions, we proved in our previous work that the admissible smooth representations $π$ of $\mathrm{GL}_2(K)$ that occur in Hecke eigenspaces of the mod $p$ cohomology are of finite length. In this paper we obtain various refined results about the structure of subquotients of $π$, such as their Iwahori-socle filtrations and $K_1$-invariants, where $K_1$ is the principal congruence subgroup of $\mathrm{GL}_2(\mathcal{O}_K)$. We also determine the Hilbert series of $π$ as Iwahori-representation under these conditions.

math.NT

To be or not to be local

Let $p$ be a prime number and $K$ a finite unramified extension of $\mathbf{Q}_p$. For a smooth representation $π$ of $\mathrm{GL}_2(K)$ occurring in some Hecke eigenspace of the mod $p$ cohomology of a Shimura curve, we explore different strategies (inspired by the case $K=\mathbf{Q}_p$) to attack the locality question: does $π$ depend only on the underlying $2$-dimensional representation $\overlineρ$ of ${\rm Gal}(\overline K/K)$? In particular when $[K:\mathbf{Q}_p]=2$, crucially using perfectoid geometry, we associate to $\overlineρ$ an infinite-dimensional mod $p$ smooth representation of $\begin{pmatrix}K^\times&K\\0&1\end{pmatrix}$ which we hope is the restriction to $\begin{pmatrix}K^\times&K\\0&1\end{pmatrix}$ of the (irreducible) supersingular subquotient of $π$.

math.NT

Finite length for unramified $\mathrm{GL}_2$

Let $p$ be a prime number and $K$ a finite unramified extension of $\mathbb{Q}_p$. If $p$ is large enough with respect to $[K:\mathbb{Q}_p]$ and under mild genericity assumptions, we prove that the admissible smooth representations of $\mathrm{GL}_2(K)$ that occur in Hecke eigenspaces of the mod $p$ cohomology are of finite length. We also prove many new structural results about these representations of $\mathrm{GL}_2(K)$ and their subquotients.

math.NT

Towards Human-AI Synergy in UI Design: Supporting Iterative Generation with LLMs

In automated UI design generation, a key challenge is the lack of support for iterative processes, as most systems focus solely on end-to-end output. This stems from limited capabilities in interpreting design intent and a lack of transparency for refining intermediate results. To better understand these challenges, we conducted a formative study that identified concrete and actionable requirements for supporting iterative design with Generative Tools. Guided by these findings, we propose PrototypeFlow, a human-centered system for automated UI generation that leverages multi-modal inputs and models. PrototypeFlow takes natural language descriptions and layout preferences as input to generate the high-fidelity UI design. At its core is a theme design module that clarifies implicit design intent through prompt enhancement and orchestrates sub-modules for component-level generation. Designers retain full control over inputs, intermediate results, and final prototypes, enabling flexible and targeted refinement by steering generation and directly editing outputs. Our experiments and user studies confirmed the effectiveness and usefulness of our proposed PrototypeFlow.

cs.HC

SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance

Speech is essential for human communication, yet millions of people face impairments such as dysarthria, stuttering, and aphasia conditions that often lead to social isolation and reduced participation. Despite recent progress in automatic speech recognition (ASR) and text-to-speech (TTS) technologies, accessible web and mobile infrastructures for users with impaired speech remain limited, hindering the practical adoption of these advances in daily communication. To bridge this gap, we present SpeechAgent, a mobile SpeechAgent designed to facilitate people with speech impairments in everyday communication. The system integrates large language model (LLM)- driven reasoning with advanced speech processing modules, providing adaptive support tailored to diverse impairment types. To ensure real-world practicality, we develop a structured deployment pipeline that enables real-time speech processing on mobile and edge devices, achieving imperceptible latency while maintaining high accuracy and speech quality. Evaluation on real-world impaired speech datasets and edge-device latency profiling confirms that SpeechAgent delivers both effective and user-friendly performance, demonstrating its feasibility for personalized, day-to-day assistive communication.

eess.SY

Actual Achieved Gain and Optimal Perceived Gain: Modeling Human Take-over Decisions Towards Automated Vehicles' Suggestions

Driver decision quality in take-overs is critical for effective human-Autonomous Driving System (ADS) collaboration. However, current research lacks detailed analysis of its variations. This paper introduces two metrics--Actual Achieved Gain (AAG) and Optimal Perceived Gain (OPG)--to assess decision quality, with OPG representing optimal decisions and AAG reflecting actual outcomes. Both are calculated as weighted averages of perceived gains and losses, influenced by ADS accuracy. Study 1 (N=315) used a 21-point Thurstone scale to measure perceived gains and losses-key components of AAG and OPG-across typical tasks: route selection, overtaking, and collision avoidance. Studies 2 (N=54) and 3 (N=54) modeled decision quality under varying ADS accuracy and decision time. Results show with sufficient time (>3.5s), AAG converges towards OPG, indicating rational decision-making, while limited time leads to intuitive and deterministic choices. Study 3 also linked AAG-OPG deviations to irrational behaviors. An intervention study (N=8) and a pilot (N=4) employing voice alarms and multi-modal alarms based on these deviations demonstrated AAG's potential to improve decision quality.

cs.HC

Adanonymizer: Interactively Navigating and Balancing the Duality of Privacy and Output Performance in Human-LLM Interaction

Current Large Language Models (LLMs) cannot support users to precisely balance privacy protection and output performance during individual consultations. We introduce Adanonymizer, an anonymization plug-in that allows users to control this balance by navigating a trade-off curve. A survey (N=221) revealed a privacy paradox, where users frequently disclosed sensitive information despite acknowledging privacy risks. The study further demonstrated that privacy risks were not significantly correlated with model output performance, highlighting the potential to navigate this trade-off. Adanonymizer normalizes privacy and utility ratings by type and automates the pseudonymization of sensitive terms based on user preferences, significantly reducing user effort. Its 2D color palette interface visualizes the privacy-utility trade-off, allowing users to adjust the balance by manipulating a point. An evaluation (N=36) compared Adanonymizer with ablation methods and differential privacy techniques, where Adanonymizer significantly reduced modification time, achieved better perceived model performance and overall user preference.

cs.HC

Exploring Device-Oriented Video Encryption for Hierarchical Privacy Protection in AR Content Sharing

Content sharing across multiple Augmented Reality (AR) displays is becoming commonplace, enhancing team communication and collaboration through devices like smartphones and AR glasses. However, this practice raises significant privacy concerns, especially concerning the physical environment visible in AR, which may include sensitive personal details like facial features and identifiable information. Our research focuses on protecting privacy within AR environments, particularly the physical backgrounds visible during content sharing across three common AR display methods: projection, smartphone, and AR glasses. We analyze the potential privacy risks associated with each method and employ a Region Of Interest (ROI) video encryption system to hierarchically encrypt the physical backdrop based on its safety rating. This study pioneers the integration of ROI video encryption at the bitstream level within AR contexts, providing a more efficient solution than traditional pixel-level encryption by enhancing encryption speed and reducing the required space. Our adaptive system dynamically adjusts the encryption intensity based on the AR display method, ensuring tailored privacy protection.

cs.HC

Multivariable ($φ$,$\mathcal{O}_K^\times$)-modules and local-global compatibility

Let $p$ be a prime number, $K$ a finite unramified extension of $\mathbb{Q}_p$ and $\mathbb{F}$ a finite extension of $\mathbb{F}_p$. Using perfectoid spaces we associate to any finite-dimensional continuous representation $\overlineρ$ of ${\rm Gal}(\overline K/K)$ over $\mathbb{F}$ an étale $(φ,\mathcal{O}_K^\times)$-module $D_A^\otimes(\overlineρ)$ over a completed localization $A$ of $\mathbb{F}[\![\mathcal{O}_K]\!]$. We conjecture that one can also associate an étale $(φ,\mathcal{O}_K^\times)$-module $D_A(π)$ to any smooth representation $π$ of $\mathrm{GL}_2(K)$ occurring in some Hecke eigenspace of the mod $p$ cohomology of a Shimura curve, and that moreover $D_A(π)$ is isomorphic (up to twist) to $D_A^\otimes(\overlineρ)$, where $\overlineρ$ is the underlying $2$-dimensional representation of ${\rm Gal}(\overline K/K)$. Using previous work of the same authors, we prove this conjecture when $\overlineρ$ is semi-simple and sufficiently generic.

math.NT

MultiSurf-GPT: Facilitating Context-Aware Reasoning with Large-Scale Language Models for Multimodal Surface Sensing

Surface sensing is widely employed in health diagnostics, manufacturing and safety monitoring. Advances in mobile sensing affords this potential for context awareness in mobile computing, typically with a single sensing modality. Emerging multimodal large-scale language models offer new opportunities. We propose MultiSurf-GPT, which utilizes the advanced capabilities of GPT-4o to process and interpret diverse modalities (radar, microscope and multispectral data) uniformly based on prompting strategies (zero-shot and few-shot prompting). We preliminarily validated our framework by using MultiSurf-GPT to identify low-level information, and to infer high-level context-aware analytics, demonstrating the capability of augmenting context-aware insights. This framework shows promise as a tool to expedite the development of more complex context-aware applications in the future, providing a faster, more cost-effective, and integrated solution.

cs.HC

Towards Enhanced Context Awareness with Vision-based Multimodal Interfaces

Vision-based Interfaces (VIs) are pivotal in advancing Human-Computer Interaction (HCI), particularly in enhancing context awareness. However, there are significant opportunities for these interfaces due to rapid advancements in multimodal Artificial Intelligence (AI), which promise a future of tight coupling between humans and intelligent systems. AI-driven VIs, when integrated with other modalities, offer a robust solution for effectively capturing and interpreting user intentions and complex environmental information, thereby facilitating seamless and efficient interactions. This PhD study explores three application cases of multimodal interfaces to augment context awareness, respectively focusing on three dimensions of visual modality: scale, depth, and time: a fine-grained analysis of physical surfaces via microscopic image, precise projection of the real world using depth data, and rendering haptic feedback from video background in virtual environments.

cs.HC

Exploring Large-Scale Language Models to Evaluate EEG-Based Multimodal Data for Mental Health

Integrating physiological signals such as electroencephalogram (EEG), with other data such as interview audio, may offer valuable multimodal insights into psychological states or neurological disorders. Recent advancements with Large Language Models (LLMs) position them as prospective ``health agents'' for mental health assessment. However, current research predominantly focus on single data modalities, presenting an opportunity to advance understanding through multimodal data. Our study aims to advance this approach by investigating multimodal data using LLMs for mental health assessment, specifically through zero-shot and few-shot prompting. Three datasets are adopted for depression and emotion classifications incorporating EEG, facial expressions, and audio (text). The results indicate that multimodal information confers substantial advantages over single modality approaches in mental health assessment. Notably, integrating EEG alongside commonly used LLM modalities such as audio and images demonstrates promising potential. Moreover, our findings reveal that 1-shot learning offers greater benefits compared to zero-shot learning methods.

cs.HC

Investigating the Design Considerations for Integrating Text-to-Image Generative AI within Augmented Reality Environments

Generative Artificial Intelligence (GenAI) has emerged as a fundamental component of intelligent interactive systems, enabling the automatic generation of multimodal media content. The continuous enhancement in the quality of Artificial Intelligence-Generated Content (AIGC), including but not limited to images and text, is forging new paradigms for its application, particularly within the domain of Augmented Reality (AR). Nevertheless, the application of GenAI within the AR design process remains opaque. This paper aims to articulate a design space encapsulating a series of criteria and a prototypical process to aid practitioners in assessing the aptness of adopting pertinent technologies. The proposed model has been formulated based on a synthesis of design insights garnered from ten experts, obtained through focus group interviews. Leveraging these initial insights, we delineate potential applications of GenAI in AR.

cs.HC

MicroCam: Leveraging Smartphone Microscope Camera for Context-Aware Contact Surface Sensing

The primary focus of this research is the discreet and subtle everyday contact interactions between mobile phones and their surrounding surfaces. Such interactions are anticipated to facilitate mobile context awareness, encompassing aspects such as dispensing medication updates, intelligently switching modes (e.g., silent mode), or initiating commands (e.g., deactivating an alarm). We introduce MicroCam, a contact-based sensing system that employs smartphone IMU data to detect the routine state of phone placement and utilizes a built-in microscope camera to capture intricate surface details. In particular, a natural dataset is collected to acquire authentic surface textures in situ for training and testing. Moreover, we optimize the deep neural network component of the algorithm, based on continual learning, to accurately discriminate between object categories (e.g., tables) and material constituents (e.g., wood). Experimental results highlight the superior accuracy, robustness and generalization of the proposed method. Lastly, we conducted a comprehensive discussion centered on our prototype, encompassing topics such as system performance and potential applications and scenarios.

cs.HC

On some mod $p$ representations of quaternion algebra over $\mathbb{Q}_p$

Let $F$ be a totally real field in which $p$ is unramified and $B$ be a quaternion algebra which splits at at most one infinite place. Let $\overline{r}:\mathrm{Gal}(\overline{F}/F)\to \mathrm{GL}_2(\overline{\mathbb{F}}_p)$ be a modular Galois representation which satisfies the Taylor-Wiles hypotheses. Assume that for some fixed place $v|p$, $B$ ramifies at $v$ and $F_v$ is isomorphic to $ \mathbb{Q}_p$ and $\overline{r}$ is generic at $v$. We prove that the admissible smooth representations of the quaternion algebra over $\mathbb{Q}_p$ coming from mod $p$ cohomology of Shimura varieties associated to $B$ have Gelfand-Kirillov dimension $1$. As an application we prove that the degree two Scholze's functor vanishes on supersingular representations of $\mathrm{GL}_2(\mathbb{Q}_p)$. We also prove some finer structure theorem about the image of Scholze's functor in the reducible case.

math.NT

AcademicGPT: Empowering Academic Research

Large Language Models (LLMs) have demonstrated exceptional capabilities across various natural language processing tasks. Yet, many of these advanced LLMs are tailored for broad, general-purpose applications. In this technical report, we introduce AcademicGPT, designed specifically to empower academic research. AcademicGPT is a continual training model derived from LLaMA2-70B. Our training corpus mainly consists of academic papers, thesis, content from some academic domain, high-quality Chinese data and others. While it may not be extensive in data scale, AcademicGPT marks our initial venture into a domain-specific GPT tailored for research area. We evaluate AcademicGPT on several established public benchmarks such as MMLU and CEval, as well as on some specialized academic benchmarks like PubMedQA, SCIEval, and our newly-created ComputerScienceQA, to demonstrate its ability from general knowledge ability, to Chinese ability, and to academic ability. Building upon AcademicGPT's foundation model, we also developed several applications catered to the academic area, including General Academic Question Answering, AI-assisted Paper Reading, Paper Review, and AI-assisted Title and Abstract Generation.

cs.CL

Conjectures and results on modular representations of $\mathrm{GL}_n(K)$ for a $p$-adic field $K$

Let $p$ be a prime number and $K$ a finite extension of $\mathbb{Q}_p$. We state conjectures on the smooth representations of $\mathrm{GL}_n(K)$ that occur in spaces of mod $p$ automorphic forms (for compact unitary groups). In particular, when $K$ is unramified, we conjecture that they are of finite length and predict their internal structure (extensions, form of subquotients) from the structure of a certain algebraic representation of $\mathrm{GL}_n$. When $n=2$ and $K$ is unramified, we prove several cases of our conjectures, including new finite length results.

math.NT

Test-takers have a say: understanding the implications of the use of AI in language tests

Language tests measure a person's ability to use a language in terms of listening, speaking, reading, or writing. Such tests play an integral role in academic, professional, and immigration domains, with entities such as educational institutions, professional accreditation bodies, and governments using them to assess candidate language proficiency. Recent advances in Artificial Intelligence (AI) and the discipline of Natural Language Processing have prompted language test providers to explore AI's potential applicability within language testing, leading to transformative activity patterns surrounding language instruction and learning. However, with concerns over AI's trustworthiness, it is imperative to understand the implications of integrating AI into language testing. This knowledge will enable stakeholders to make well-informed decisions, thus safeguarding community well-being and testing integrity. To understand the concerns and effects of AI usage in language tests, we conducted interviews and surveys with English test-takers. To the best of our knowledge, this is the first empirical study aimed at identifying the implications of AI adoption in language tests from a test-taker perspective. Our study reveals test-taker perceptions and behavioral patterns. Specifically, we identify that AI integration may enhance perceptions of fairness, consistency, and availability. Conversely, it might incite mistrust regarding reliability and interactivity aspects, subsequently influencing the behaviors and well-being of test-takers. These insights provide a better understanding of potential societal implications and assist stakeholders in making informed decisions concerning AI usage in language testing.

cs.CY