SearcharxivSearch

arXiv subjects

Jiangang Hao

Publications and source records attributed to Jiangang Hao.

At least 19 recordsLinked to original sources

Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.

cs.CL

Automated Coding of Communication Data Using ChatGPT: Consistency Across Subgroups

Assessing communication and collaboration at scale depends on a labor-intensive task of coding communication data into categories according to different frameworks. Prior research has established that ChatGPT can be directly instructed with coding rubrics to code the communication data and achieves accuracy comparable to human raters. However, whether the coding from ChatGPT or similar AI technology perform consistently across different demographic groups, such as gender and race, remains unclear. To address this gap, we introduce three checks for evaluating subgroup consistency in LLM-based coding by adapting an existing framework from the automated scoring literature. Using a typical collaborative problem-solving coding framework and data from three types of collaborative tasks, we examine ChatGPT-based coding performance across gender and racial/ethnic groups. Our results show that ChatGPT-based coding perform consistently in the same way as human raters across gender or racial/ethnic groups, demonstrating the possibility of its use in large-scale assessments of collaboration and communication.

cs.CL

Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs

Writing is a foundational literacy skill that underpins effective communication, fosters critical thinking, facilitates learning across disciplines, and enables individuals to organize and articulate complex ideas. Consequently, writing assessment plays a vital role in evaluating language proficiency, communicative effectiveness, and analytical reasoning. The rapid advancement of large language models (LLMs) has made it increasingly easy to generate coherent, high-quality essays, raising significant concerns about the authenticity of student-submitted work. This chapter first provides an overview of the current landscape of detectors for AI-generated and AI-assisted essays, along with guidelines for their responsible use. It then presents empirical analyses to evaluate how well detectors trained on essays from one LLM generalize to identifying essays produced by other LLMs, based on essays generated in response to public GRE writing prompts. These findings provide guidance for developing and retraining detectors for practical applications.

cs.CL

AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity

The rapid advancement of large language models (LLMs) has enabled the generation of coherent essays, making AI-assisted writing increasingly common in educational and professional settings. Using large-scale empirical data, we examine and benchmark the characteristics and quality of essays generated by popular LLMs and discuss their implications for two key components of writing assessments: automated scoring and academic integrity. Our findings highlight limitations in existing automated scoring systems, such as e-rater, when applied to essays generated or heavily influenced by AI, and identify areas for improvement, including the development of new features to capture deeper thinking and recalibrating feature weights. Despite growing concerns that the increasing variety of LLMs may undermine the feasibility of detecting AI-generated essays, our results show that detectors trained on essays generated from one model can often identify texts from others with high accuracy, suggesting that effective detection could remain manageable in practice.

cs.CL

Automated Coding of Communications in Collaborative Problem-solving Tasks Using ChatGPT

Collaborative problem solving (CPS) is widely recognized as a critical 21st-century skill. Assessing CPS depends heavily on coding the communication data using a construct-relevant framework, and this process has long been a major bottleneck to scaling up such assessments. Based on five datasets and two coding frameworks, we demonstrate that ChatGPT can code communication data to a satisfactory level, though performance varies across ChatGPT models, and depends on the coding framework and task characteristics. Interestingly, newer reasoning-focused models such as GPT-o1-mini and GPT-o3-mini do not necessarily yield better coding results. Additionally, we show that refining prompts based on feedback from miscoded cases can improve coding accuracy in some instances, though the effectiveness of this approach is not consistent across all tasks. These findings offer practical guidance for researchers and practitioners in developing scalable, efficient methods to analyze communication data in support of 21st-century skill assessment.

cs.HC

Test Security in Remote Testing Age: Perspectives from Process Data Analytics and AI

The COVID-19 pandemic has accelerated the implementation and acceptance of remotely proctored high-stake assessments. While the flexible administration of the tests brings forth many values, it raises test security-related concerns. Meanwhile, artificial intelligence (AI) has witnessed tremendous advances in the last five years. Many AI tools (such as the very recent ChatGPT) can generate high-quality responses to test items. These new developments require test security research beyond the statistical analysis of scores and response time. Data analytics and AI methods based on clickstream process data can get us deeper insight into the test-taking process and hold great promise for securing remotely administered high-stakes tests. This chapter uses real-world examples to show that this is indeed the case.

cs.CR

Galaxy clusters in the SDSS Stripe 82 based on galaxy photometric redshifts

Based on a recent photometric redshift galaxy catalogue, we have searched for galaxy clusters in the Stripe~82 region of the Sloan Digital Sky Survey by applying the Adami & MAzure Cluster FInder (AMACFI). Extensive tests were made to fine-tune the AMACFI parameters and make the cluster detection as reliable as possible. The same method was applied to the Millennium simulation to estimate our detection efficiency and the approximate masses of the detected clusters. Considering all the cluster galaxies (i.e. within a 1 Mpc radius of the cluster to which they belong and with a photoz differing by less than 0.05 from that of the cluster), we stacked clusters in various redshift bins to derive colour-magnitude diagrams and galaxy luminosity functions (GLFs). For each galaxy with absolute magnitude brighter than -19.0 in the r band, we computed the disk and spheroid components by applying SExtractor, and by stacking clusters we determined how the disk-to-spheroid flux ratio varies with cluster redshift and mass. We detected 3663 clusters in the redshift range 0.15<z<0.70, with estimated mean masses between 10^13 and a few 10^{14 solar masses. By stacking the cluster galaxies in various redshift bins, we find a clear red sequence in the (g'-r') versus r' colour-magnitude diagrams, and the GLFs are typical of clusters, though with a possible contamination from field galaxies. The morphological analysis of the cluster galaxies shows that the fraction of late-type to early-type galaxies shows an increase with redshift (particularly in high mass clusters) and a decrease with detection level, i.e. cluster mass. From the properties of the cluster galaxies, the majority of the candidate clusters detected here seem to be real clusters with typical cluster properties.

astro-ph.CO

Orientation Bias of Optically Selected Galaxy Clusters and Its Impact on Stacked Weak Lensing Analyses

Weak-lensing measurements of the averaged shear profiles of galaxy clusters binned by some proxy for cluster mass are commonly converted to cluster mass estimates under the assumption that these cluster stacks have spherical symmetry. In this paper we test whether this assumption holds for optically selected clusters binned by estimated optical richness. Using mock catalogues created from N-body simulations populated realistically with galaxies, we ran a suite of optical cluster finders and estimated their optical richness. We binned galaxy clusters by true cluster mass and estimated optical richness and measure the ellipticity of these stacks. We find that the processes of optical cluster selection and richness estimation are biased, leading to stacked structures that are elongated along the line-of-sight. We show that weak-lensing alone cannot measure the size of this orientation bias. Weak lensing masses of stacked optically selected clusters are overestimated by up to 3-6 per cent when clusters can be uniquely associated with haloes. This effect is large enough to lead to significant biases in the cosmological parameters derived from large surveys like the Dark Energy Survey, if not calibrated via simulations or fitted simultaneously. This bias probably also contributes to the observed discrepancy between the observed and predicted Sunyaev-Zel'dovich signal of optically-selected clusters.

astro-ph.CO

The SOAR Gravitational Arc Survey - I: Survey overview and photometric catalogs

We present the first results of the SOAR (Southern Astrophysical Research) Gravitational Arc Survey (SOGRAS). The survey imaged 47 clusters in two redshift intervals centered at $z=0.27$ and $z=0.55$, targeting the richest clusters in each interval. Images were obtained in the $g'$, $r'$ and $i'$ bands using the SOAR Optical Imager (SOI), with a median seeing of 0.83, 0.76 and 0.71 arcsec, respectively, in these filters. Most of the survey clusters are located within the Sloan Digital Sky Survey (SDSS) Stripe 82 region and all of them are in the SDSS footprint. Photometric calibration was therefore performed using SDSS stars located in our SOI fields. We reached for galaxies in all fields the detection limits of $g \sim 23.5$, $r \sim 23$ and $i \sim 22.5$ for a signal-to-noise ratio (S/N) = 3. As a by-product of the image processing, we generated a source catalogue with 19760 entries, the vast majority of which are galaxies, where we list their positions, magnitudes and shape parameters. We compared our galaxy shape measurements to those of local galaxies and concluded that they were not strongly affected by seeing. From the catalogue data, we are able to identify a red sequence of galaxies in most clusters in the lower $z$ range. We found 16 gravitational arc candidates around 8 clusters in our sample. They tend to be bluer than the central galaxies in the lensing cluster. A preliminary analysis indicates that $\sim 10%$ of the clusters have arcs around them, with a possible indication of a larger efficiency associated to the high-$z$ systems when compared to the low-$z$ ones. Deeper follow-up images with Gemini strengthen the case for the strong lensing nature of the candidates found in this survey.

astro-ph.CO

The Universal Einstein Radius Distribution from 10,000 SDSS Clusters

We present results from strong-lens modelling of 10,000 SDSS clusters, to establish the universal distribution of Einstein radii. Detailed lensing analyses have shown that the inner mass distribution of clusters can be accurately modelled by assuming light traces mass, successfully uncovering large numbers of multiple-images. Approximate critical curves and the effective Einstein radius of each cluster can therefore be readily calculated, from the distribution of member galaxies and scaled by their luminosities. We use a subsample of 10 well-studied clusters covered by both SDSS and HST to calibrate and test this method, and show that an accurate determination of the Einstein radius and mass can be achieved by this approach "blindly", in an automated way, and without requiring multiple images as input. We present the results of the first 10,000 clusters analysed in the range $0.1 =0.73^{+0.02}_{-0.03}$, $σ=0.316^{+0.004}_{-0.002}$, and with higher abundance of large $θ_{e}$ clusters than predicted by $Λ$CDM. We visually inspect each of the clusters with $θ_{e}>40 \arcsec$ ($z_{s}=2$) and find that $\sim20%$ are boosted by various projection effects detailed here, remaining with $\sim40$ real giant-lens candidates, with a maximum of $θ_{e}=69\pm12 \arcsec$ ($z_{s}=2$) for the most massive candidate, in agreement with semi-analytic calculations. The results of this work should be verified further when an extended calibration sample is available.

astro-ph.CO

The SDSS Coadd: Cosmic Shear Measurement

Stripe 82 in the Sloan Digital Sky Survey was observed multiple times, allowing deeper images to be constructed by coadding the data. Here we analyze the ellipticities of background galaxies in this 275 square degree region, searching for evidence of distortions due to cosmic shear. The E-mode is detected in both real and Fourier space with $>5$-$σ$ significance on degree scales, while the B-mode is consistent with zero as expected. The amplitude of the signal constrains the combination of the matter density $Ω_m$ and fluctuation amplitude $σ_8$ to be $Ω_m^{0.7}σ_8 = 0.252^{+0.032}_{-0.052}$.

astro-ph.CO

The SDSS Coadd: A Galaxy Photometric Redshift Catalog

We present and describe a catalog of galaxy photometric redshifts (photo-z's) for the Sloan Digital Sky Survey (SDSS) Coadd Data. We use the Artificial Neural Network (ANN) technique to calculate photo-z's and the Nearest Neighbor Error (NNE) method to estimate photo-z errors for $\sim$ 13 million objects classified as galaxies in the coadd with $r < 24.5$. The photo-z and photo-z error estimators are trained and validated on a sample of $\sim 83,000$ galaxies that have SDSS photometry and spectroscopic redshifts measured by the SDSS Data Release 7 (DR7), the Canadian Network for Observational Cosmology Field Galaxy Survey (CNOC2), the Deep Extragalactic Evolutionary Probe Data Release 3(DEEP2 DR3), the VIsible imaging Multi-Object Spectrograph - Very Large Telescope Deep Survey (VVDS) and the WiggleZ Dark Energy Survey. For the best ANN methods we have tried, we find that 68% of the galaxies in the validation set have a photo-z error smaller than $σ_{68} =0.031$. After presenting our results and quality tests, we provide a short guide for users accessing the public data.

astro-ph.CO

The SDSS Coadd: Cross-Correlation Weak Lensing and Tomography of Galaxy Clusters

The shapes of distant galaxies are sheared by intervening galaxy clusters. We examine this effect in Stripe 82, a 275 square degree region observed multiple times in the Sloan Digital Sky Survey and coadded to achieve greater depth. We obtain a mass-richness calibration that is similar to other SDSS analyses, demonstrating that the coaddition process did not adversely affect the lensing signal. We also propose a new parameterization of the effect of tomography on the cluster lensing signal which does not require binning in redshift, and we show that using this parameterization we can detect tomography for stacked clusters at varying redshifts. Finally, due to the sensitivity of the tomographic detection to accurately marginalizing over the effect of the cluster mass, we show that tomography at low redshift (where dependence on exact cosmological models is weak) can be used to constrain mass profiles in clusters.

astro-ph.CO

The SDSS Coadd: 275 deg^2 of Deep SDSS Imaging on Stripe 82

We present details of the construction and characterization of the coaddition of the Sloan Digital Sky Survey Stripe 82 \ugriz\ imaging data. This survey consists of 275 deg$^2$ of repeated scanning by the SDSS camera of $2.5\arcdeg$ of $δ$ over $-50\arcdeg \le α\le 60\arcdeg$ centered on the Celestial Equator. Each piece of sky has $\sim 20$ runs contributing and thus reaches $\sim2$ magnitudes fainter than the SDSS single pass data, i.e. to $r\sim 23.5$ for galaxies. We discuss the image processing of the coaddition, the modeling of the PSF, the calibration, and the production of standard SDSS catalogs. The data have $r$-band median seeing of 1.1\arcsec, and are calibrated to $\le 1%$. Star color-color, number counts, and psf size vs modelled size plots show the modelling of the PSF is good enough for precision 5-band photometry. Structure in the psf-model vs magnitude plot show minor psf mis-modelling that leads to a region where stars are being mis-classified as galaxies, and this is verified using VVDS spectroscopy. As this is a wide area deep survey there are a variety of uses for the data, including galactic structure, photometric redshift computation, cluster finding and cross wavelength measurements, weak lensing cluster mass calibrations, and cosmic shear measurements.

astro-ph.CO

Intrinsic Alignment of Cluster Galaxies: the Redshift Evolution

We present measurements of two types of cluster galaxy alignments based on a volume limited and highly pure ($\ge$ 90%) sample of clusters from the GMBCG catalog derived from SDSS DR7. We detect a clear BCG alignment (the alignment of major axis of the BCG toward the distribution of cluster satellite galaxies). We find that the BCG alignment signal becomes stronger as the redshift and BCG absolute magnitude decrease, and becomes weaker as BCG stellar mass decreases. No dependence of the BCG alignment on cluster richness is found. We can detect a statistically significant ($\ge$ 3 sigma) satellite alignment (the alignment of the major axes of the cluster satellite galaxies toward the BCG) only when we use the isophotal fit position angles (PAs, hereafter), and the satellite alignment depends on the apparent magnitudes rather than the absolute magnitudes of the BCGs. This suggests the detected satellite alignment based on isophotoal PAs from the SDSS pipeline is possibly due to the contamination from the diffuse light of nearby BCGs. We caution that this should not be simply interpreted as non-existence of the satellite alignment, but rather that we cannot detect them with our current photometric SDSS data. We perform our measurements on both SDSS $r$ band and $i$ band data, but did not observe a passband dependence of the alignments.

astro-ph.CO

The Sunyaev-Zeldovich Signal of the maxBCG SDSS Galaxy Clusters in WMAP

The Planck Collaboration measured the Sunyaev-Zel'dovich (SZ) decrement of optically selected clusters from the Sloan Digital Sky Survey, finding that it falls significantly below expectations based on existing mass calibration of the maxBCG galaxy clusters. Resolving this tension requires either the data to go up, or the theoretical expectations to come down. Here, we use data from the Wilkinson Microwave Anisotropy Probe (WMAP) to perform an independent estimate of the SZ decrement of maxBCG clusters. The recovered signal is consistent with that obtained using Planck, though with larger error bars due to WMAP's larger beam size and smaller frequency range. Nevertheless, this detection serves as an independent confirmation of the magnitude of the effect, and demonstrates that the observed discrepancy must be theoretical in origin.

astro-ph.CO

A GMBCG Galaxy Cluster Catalog of 55,424 Rich Clusters from SDSS DR7

We present a large catalog of optically selected galaxy clusters from the application of a new Gaussian Mixture Brightest Cluster Galaxy (GMBCG) algorithm to SDSS Data Release 7 data. The algorithm detects clusters by identifying the red sequence plus Brightest Cluster Galaxy (BCG) feature, which is unique for galaxy clusters and does not exist among field galaxies. Red sequence clustering in color space is detected using an Error Corrected Gaussian Mixture Model. We run GMBCG on 8240 square degrees of photometric data from SDSS DR7 to assemble the largest ever optical galaxy cluster catalog, consisting of over 55,000 rich clusters across the redshift range from 0.1 < z < 0.55. We present Monte Carlo tests of completeness and purity and perform cross-matching with X-ray clusters and with the maxBCG sample at low redshift. These tests indicate high completeness and purity across the full redshift range for clusters with 15 or more members.

astro-ph.CO

The Sloan Bright Arcs Survey : Discovery of Seven New Strongly Lensed Galaxies from z=0.66-2.94

We report the discovery of seven new, very bright gravitational lens systems from our ongoing gravitational lens search, the Sloan Bright Arcs Survey (SBAS). Two of the systems are confirmed to have high source redshifts z=2.19 and z=2.94. Three other systems lie at intermediate redshift with z=1.33,1.82,1.93 and two systems are at low redshift z=0.66,0.86. The lensed source galaxies in all of these systems are bright, with i-band magnitudes ranging from 19.73-22.06. We present the spectrum of each of the source galaxies in these systems along with estimates of the Einstein radius for each system. The foreground lens in most systems is identified by a red sequence based cluster finder as a galaxy group; one system is identified as a moderately rich cluster. In total the SBAS has now discovered 19 strong lens systems in the SDSS imaging data, 8 of which are among the highest surface brightness z\simeq2-3 galaxies known.

astro-ph.CO