SearcharxivSearch

arXiv subjects

Jiyoung Han

Publications and source records attributed to Jiyoung Han.

At least 19 recordsLinked to original sources

Visual Framing for News Stance Detection via Image Generation

Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct challenges because their stances are often implicit, subtly conveyed through journalistic framing, and embedded in long, structurally complex texts. To address these challenges, we introduce VFStance, which leverages visual framing to make implicit stance cues more explicit via image generation. In evaluation experiments, we demonstrate the effectiveness of VFStance over existing methods and the contribution of visual framing to its performance. Finally, a controlled user study (N=200) in a snippet-based news consumption setting further demonstrates that VFStance can make stance signals visually salient and highlights its potential use beyond automated stance detection.

cs.CL

Quantitative Oppenheim Conjecture for Random Quadratic Forms and Optimal Variance Bounds in Function Fields

We prove a quantitative version of Oppenheim's conjecture in the function field setting. In order to do so, we compute the higher moments of the Siegel transform. In particular, we find an optimal bound on the variance of the number of lattice points in a set. Moreover, we compute the exact variance of the number of lattice points in a ball, which is of independent interest.

math.NT

Distribution of values at tuples of integer vectors under symplectic forms

We investigate lattice-counting problems associated with symplectic forms from the perspective of homogeneous dynamics. In the qualitative direction, we establish an analog of Margulis theorem for symplectic forms, proving density results for tuples of vectors. Quantitatively, we derive a volume formula having a certain growth rate, and use this and Rogers' formulas for a higher rank Siegel transform to obtain the asymptotic formulas of the counting function associated with a generic symplectic form. We further establish primitive and congruent analogs of the generic quantitative result. For the primitive case, we show that the lack of completely explicit higher moment formulas for a primitive higher rank Siegel transform does not obstruct obtaining quantitative statements.

math.DS

Journalism-Guided Agentic In-Context Learning for News Stance Detection

As online news consumption grows, personalized recommendation systems have become integral to digital journalism. However, these systems risk reinforcing filter bubbles and political polarization by failing to incorporate diverse perspectives. Stance detection -- identifying a text's position on a target -- can help mitigate this by enabling viewpoint-aware recommendations and data-driven analyses of media bias. Yet, existing stance detection research remains largely limited to short texts and high-resource languages. To address these gaps, we introduce \textsc{K-News-Stance}, the first Korean dataset for article-level stance detection, comprising 2,000 news articles with article-level and 21,650 segment-level stance annotations across 47 societal issues. We also propose \textsc{JoA-ICL}, a \textbf{Jo}urnalism-guided \textbf{A}gentic \textbf{I}n-\textbf{C}ontext \textbf{L}earning framework that employs a language model agent to predict the stances of key structural segments (e.g., leads, quotations), which are then aggregated to infer the overall article stance. Experiments showed that \textsc{JoA-ICL} outperforms existing stance detection methods, highlighting the benefits of segment-level agency in capturing the overall position of long-form news articles. Two case studies further demonstrate its broader utility in promoting viewpoint diversity in news recommendations and uncovering patterns of media bias.

cs.CL

Moment formulas of Siegel transforms with congruence conditions in dimension 2

We compute the first and second moment formulas for Siegel transforms related to problems counting primitive lattice points in the real plane with congruence conditions. As applications, we derive an analog of Schmidt's random counting theorem and the quantitative Khintchine theorem for irrational numbers, approximated by rational numbers $p/q$, where we place a congruence-conditional constraint on the vector $(p,q)$.

math.NT

Deep literature reviews: an application of fine-tuned language models to migration research

This paper presents a hybrid framework for literature reviews that augments traditional bibliometric methods with large language models (LLMs). By fine-tuning open-source LLMs, our approach enables scalable extraction of qualitative insights from large volumes of research content, enhancing both the breadth and depth of knowledge synthesis. To improve annotation efficiency and consistency, we introduce an error-focused validation process in which LLMs generate initial labels and human reviewers correct misclassifications. Applying this framework to over 20000 scientific articles about human migration, we demonstrate that a domain-adapted LLM can serve as a "specialist" model - capable of accurately selecting relevant studies, detecting emerging trends, and identifying critical research gaps. Notably, the LLM-assisted review reveals a growing scholarly interest in climate-induced migration. However, existing literature disproportionately centers on a narrow set of environmental hazards (e.g., floods, droughts, sea-level rise, and land degradation), while overlooking others that more directly affect human health and well-being, such as air and water pollution or infectious diseases. This imbalance highlights the need for more comprehensive research that goes beyond physical environmental changes to examine their ecological and societal consequences, particularly in shaping migration as an adaptive response. Overall, our proposed framework demonstrates the potential of fine-tuned LLMs to conduct more efficient, consistent, and insightful literature reviews across disciplines, ultimately accelerating knowledge synthesis and scientific discovery.

cs.CL

Persona Setting Pitfall: Persistent Outgroup Biases in Large Language Models Arising from Social Identity Adoption

Drawing parallels between human cognition and artificial intelligence, we explored how large language models (LLMs) internalize identities imposed by targeted prompts. Informed by Social Identity Theory, these identity assignments lead LLMs to distinguish between "we" (the ingroup) and "they" (the outgroup). This self-categorization generates both ingroup favoritism and outgroup bias. Nonetheless, existing literature has predominantly focused on ingroup favoritism, often overlooking outgroup bias, which is a fundamental source of intergroup prejudice and discrimination. Our experiment addresses this gap by demonstrating that outgroup bias manifests as strongly as ingroup favoritism. Furthermore, we successfully mitigated the inherent pro-liberal, anti-conservative bias in LLMs by guiding them to adopt the perspectives of the initially disfavored group. These results were replicated in the context of gender bias. Our findings highlight the potential to develop more equitable and balanced language models.

cs.CL

Distribution of Primitive Lattice Points in Large Dimensions

We investigate the asymptotic behavior of the distribution of primitive lattice points in a symmetric Borel set $S_d\subset\mathbb R^d$ as $d$ goes to infinity, under certain volume conditions on $S_d$. Our main technique involves exploring higher moment formulas for the primitive Siegel transform. We first demonstrate that if the volume of $S_d$ remains fixed for all $d\in \mathbb N$, then the distribution of the half the number of primitive lattice points in $S_d$ converges, in distribution, to the Poisson distribution of mean $\frac 1 2$. Furthermore, if the volume of $S_d$ goes to infinity subexponentially as $d$ approaches infinity, the normalized distribution of the half the number of primitive lattice points in $S_d$ converges, in distribution, to the normal distribution $\mathcal N(0,1)$. We also extend these results to the setting of stochastic processes. This work is motivated by the contributions of Rogers (1955), S\"odergren (2011) and Str\"ombergsson and S\"odergren (2019).

math.NT

Weight decomposition of $\mathfrak{sl}_d(\mathbb R)$ with respect to the adjoint representation of $\mathfrak{so}(p,q)$

In this concise article, we compute the weight decomposition of $\mathfrak{sl}_d(\mathbb R)$ with respect to the adjoint representation of $\mathfrak{so}(p,q)$, where $d=p+q$ and demonstrate in detail that $\mathfrak{sl}_d(\mathbb R)$ comprises two irreducible $\mathfrak{so}(p,q)$-invariant subspaces. This can be employed to establish the well-known fact that the identity component of $\mathrm{SO}(p,q)$ is a maximal connected subgroup of $\mathrm{SL}_d(\mathbb R)$.

math.RT

I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models

We explored cultural biases-individualism vs. collectivism-in ChatGPT across three Western languages (i.e., English, German, and French) and three Eastern languages (i.e., Chinese, Japanese, and Korean). When ChatGPT adopted an individualistic persona in Western languages, its collectivism scores (i.e., out-group values) exhibited a more negative trend, surpassing their positive orientation towards individualism (i.e., in-group values). Conversely, when a collectivistic persona was assigned to ChatGPT in Eastern languages, a similar pattern emerged with more negative responses toward individualism (i.e., out-group values) as compared to collectivism (i.e., in-group values). The results indicate that when imbued with a particular social identity, ChatGPT discerns in-group and out-group, embracing in-group values while eschewing out-group values. Notably, the negativity towards the out-group, from which prejudices and discrimination arise, exceeded the positivity towards the in-group. The experiment was replicated in the political domain, and the results remained consistent. Furthermore, this replication unveiled an intrinsic Democratic bias in Large Language Models (LLMs), aligning with earlier findings and providing integral insights into mitigating such bias through prompt engineering. Extensive robustness checks were performed using varying hyperparameter and persona setup methods, with or without social identity labels, across other popular language models.

cs.CL

Mean value theorems for the S-arithmetic primitive Siegel transforms

We develop the theory and properties of primitive unimodular $S$-arithmetic lattices in $\mathbb{Q}_S^d$ by giving integral formulas in the spirit of Siegel's primitive mean value formula and Rogers' and Schmidt's second moment formulas. When $d=2$, unlike in the real case, functions arising from the $S$-primitive Siegel transform are unbounded, requiring a careful analysis to establish their integrability. We then use mean value and second moment formulas in three applications. First, we obtain quantitative estimates for counting primitive $S$-arithmetic lattice points. We next establish a quantitative Khintchine--Groshev theorem, which, in the real case, involves counting primitive integer points in $\mathbb{Z}^d$ subject to congruence conditions. Finally, we derive an $S$-arithmetic logarithm law for unipotent flows in the spirit of Athreya--Margulis. These applications follow the spirit of the real case, but require new technical aspects of the proofs, particularly when $d=2$.

math.NT

Disentangling Structure and Style: Political Bias Detection in News by Inducing Document Hierarchy

We address an important gap in detecting political bias in news articles. Previous works that perform document classification can be influenced by the writing style of each news outlet, leading to overfitting and limited generalizability. Our approach overcomes this limitation by considering both the sentence-level semantics and the document-level rhetorical structure, resulting in a more robust and style-agnostic approach to detecting political bias in news articles. We introduce a novel multi-head hierarchical attention model that effectively encodes the structure of long documents through a diverse ensemble of attention heads. While journalism follows a formalized rhetorical structure, the writing style may vary by news outlet. We demonstrate that our method overcomes this domain dependency and outperforms previous approaches for robustness and accuracy. Further analysis and human evaluation demonstrate the ability of our model to capture common discourse structures in journalism. Our code is available at: https://github.com/xfactlab/emnlp2023-Document-Hierarchy

cs.CL

Detecting Contextomized Quotes in News Headlines by Contrastive Learning

Quotes are critical for establishing credibility in news articles. A direct quote enclosed in quotation marks has a strong visual appeal and is a sign of a reliable citation. Unfortunately, this journalistic practice is not strictly followed, and a quote in the headline is often "contextomized." Such a quote uses words out of context in a way that alters the speaker's intention so that there is no semantically matching quote in the body text. We present QuoteCSE, a contrastive learning framework that represents the embedding of news quotes based on domain-driven positive and negative samples to identify such an editorial strategy. The dataset and code are available at https://github.com/ssu-humane/contextomized-quote-contrastive.

cs.CL

A quantitative Khintchine-Groshev theorem for S-arithmetic Diophantine approximation

In his 1960 paper, Schmidt studied a quantitative type of Khintchine-Groshev theorem for general (higher) dimensions. Recently, a new proof of the theorem was found, which made it possible to relax the dimensional constraint and more generally, to add on the congruence condition by M. Alam, A. Ghosh, and S. Yu. In this paper, we generalize this new approach to S-arithmetic spaces and obtain a quantitative version of an S-arithmetic Khintchine-Groshev theorem. In fact, we consider a new S-arithmetic analog of Diophantine approximation, which is different from the one formerly established (see the 2007 paper of D. Kleinbock and G. Tomanov). Hence for the sake of completeness, we also deal with the convergence case of the Khintchine-Groshev theorem, based on this new generalization.

math.NT

Asymptotic distribution for pairs of linear and quadratic forms at integral vectors

We study the joint distribution of values of a pair consisting of a quadratic form $q$ and a linear form $\mathbf l$ over the set of integral vectors, a problem initiated by Dani-Margulis (1989). In the spirit of the celebrated theorem of Eskin, Margulis and Mozes on the quantitative version of the Oppenheim conjecture, we show that if $n \ge 5$ then under the assumptions that for every $(\alpha, \beta ) \in \mathbb R^2 \setminus \{ (0,0) \}$, the form $\alpha q + \beta \mathbf l^2$ is irrational and that the signature of the restriction of $q$ to the kernel of $\mathbf l$ is $(p, n-1-p)$, where $3\le p \le n-2$, the number of vectors $v \in \mathbb Z^n$ for which $\|v\| < T$, $a < q(v) < b$ and $c< \mathbf l(v) < d$ is asymptotically $$ C(q, \mathbf l)(d-c)(b-a)T^{n-3} , $$ as $T \to \infty$, where $C(q, \mathbf l)$ only depends on $q$ and $\mathbf l$. The density of the set of joint values of $(q, \mathbf l)$ under the same assumptions is shown by Gorodnik (2004).

math.DS

Higher moment formulae and limiting distributions of lattice points

We prove functional limit theorems for lattice point counting for affine and congruence lattices using the method of moments. Our main tools are higher moment formulae for Siegel transforms on the corresponding homogeneous spaces, which we believe to be of independent interest.

math.NT

Values of Inhomogeneous Forms at S-integral points

We prove effective versions of Oppenheim's conjecture for generic inhomogeneous forms in the S-arithmetic setting. We prove an effective result for fixed rational shifts and generic forms and we also prove a result where both the quadratic form and the shift are allowed to vary. In order to do so, we prove analogues of Rogers' moment formulae for $S$-arithmetic congruence quotients as well as for the space of affine lattices. We believe the latter results to be of independent interest.

math.DS

Risk Communication in Asian Countries: COVID-19 Discourse on Twitter

COVID-19 has become one of the most widely talked about topics on social media. This research characterizes risk communication patterns by analyzing the public discourse on the novel coronavirus from four Asian countries: South Korea, Iran, Vietnam, and India, which suffered the outbreak to different degrees. The temporal analysis shows that the official epidemic phases issued by governments do not match well with the online attention on COVID-19. This finding calls for a need to analyze the public discourse by new measures, such as topical dynamics. Here, we propose an automatic method to detect topical phase transitions and compare similarities in major topics across these countries over time. We examine the time lag difference between social media attention and confirmed patient counts. For dynamics, we find an inverse relationship between the tweet count and topical diversity.

cs.SI