SearcharxivSearch

arXiv subjects

Vincent Ginis

Publications and source records attributed to Vincent Ginis.

At least 19 recordsLinked to original sources

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluate paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions. We test a numeric/concision prompt that announces a token limit for Qwen3-14B and the trained effort settings of gpt-oss-20b and -120b. The Qwen prompt shortens reasoning traces by 12-17%, while accuracy changes at matched token limits are small and mixed. A concise/early-answer instruction raises MMLU-Pro accuracy by 3.8 percentage points at 512 tokens, including +2.7 points when both runs are unfinished. Its gain at 2,048 tokens is uncertain. For gpt-oss, candidate-logit answers from completed low- and medium-effort reasoning are 14.5-26.3 points more accurate than matched-horizon high-effort answers. Most of the 512-token advantage comes from lower effort finishing earlier, while differences among unfinished runs are smaller and mixed. Wrong early answers often concentrate probability on the chosen option, so earlier stopping does not uniformly improve probability quality. In these tests, a tight deadline can favor lower effort or a concise instruction, whereas allowing high effort to finish can recover higher final accuracy. Evaluations should report correct completion before a deadline, the answer obtained when a run is stopped, differences among unfinished runs, and probability assigned to the correct answer separately.

cs.LG

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds each completed run fixed and varies only what a reader inspects to judge whether the answer is correct: the response, a self-summary the model writes from the trace, the trace itself, and internal signals, each with and without the prompt. Across three benchmarks and five open-weight Qwen3 and gpt-oss models, we train matched linear correctness predictors on each access level. Without the prompt, summaries carry most of the trace's ranking signal (mean AUROC 0.774 versus 0.813) and add +0.156 over the response alone. With the prompt visible, the summary's gain collapses to +0.019, while the trace still adds +0.041. Even at equal length, the trace's last words predict correctness as well as summaries, or slightly better, and carry denser and more discriminative uncertainty and self-correction cues. On MMLU-Pro questions with both correct and incorrect runs, linear summary readers are near chance and trace readers retain only modest signal, both with and without the prompt (prompt-withheld AUROC 0.503-0.545 versus 0.544-0.590). With the prompt withheld, a GPT-5-mini reader recovers substantially more signal from both summaries and traces on gpt-oss-20b, and even then the trace keeps a small +0.034 advantage. Much of the linear readers' trace signal is associated with length. In the common case where users already hold the prompt, summaries are less helpful than the full trace for monitoring correctness. Monitorability is thus a joint property of the display and the reader, so any monitorability claim, including for faithfulness, should specify both.

cs.LG

Geometric Configurations of Perturbed Jailbreak Prompts

Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "\.C\.C" in the 1$ Llama model, display a significant association with a compliant-labeled answer.

cs.CR

The Type III realisation conjecture of Kirkland and \v{S}migoc

Kirkland and \v{S}migoc constructed a family of stochastic matrices realising the Type III boundary polynomials in the Karpelevi\v{c} region and conjectured that, conversely, every stochastic realisation of such a polynomial must come from their construction. We prove this conjecture for the full nonzero parameter range $0<\alpha\le1$, for genuine Type III reduced Ito polynomials of order $n$, $f_\alpha(x)=x^y(x^q-(1-\alpha))^d-\alpha^d$, where $n=qd+y$. For $0<\alpha<1$, the proof first reduces every realisation to a two-shift cyclic normal form using the Dmitriev--Dynkin boundary theorem. The remaining argument is finite and combinatorial: Coates' coefficient formula and the equality case of a weighted Tur\'an theorem force the $q$-cycles associated with the backward edges to split into $d$ complete multipartite classes of equal total weight. A circular-arc telescoping argument then converts this additive equality into the product condition required by Kirkland and \v{S}migoc. The endpoint $\alpha=1$ is treated separately. We also explain why the closed endpoint $\alpha=0$ is degenerate: the literal extension to this endpoint fails, because reducible realisations with closed $q$-cycles and transient states need not contain the global $n$-cycle.

math.RA

Neural Networks for Inverse Design of Cascaded-Mode Near-Field Landscapes

Structuring optical near-fields is important for applications in microscopy and nanoparticle manipulation. Traditionally, near-fields are structured using antenna nanostructures that locally convert propagating far-fields into bound near-fields. Recently, a remote structuring approach was proposed using cascaded mode interference in a multimode waveguide. However, determining the complex coefficients of the optimal modal combination needed to obtain specific near-fields remains a challenge. We address this inverse design problem using artificial neural networks. We model the relationship between the design parameters and near-field landscapes using multilayer neural networks. After training, these networks are used for gradient-based optimization to reconstruct target near-field profiles. We implement this methodology to design longitudinal and lateral field variations. Our approach designs simple and complex longitudinal landscapes, demonstrating accurate prediction and flexibility. Lateral field reconstruction is more challenging but improved with training data selection and augmentation. This work establishes deep learning as an efficient and scalable framework for cascaded-mode near-field inverse design.

physics.optics

Rebuttals Move Peer-Review Scores, but Initial-Review Structure Bounds the Movement

Author rebuttals are the main post-submission window in peer review, but their effect on reviewer scores remains hard to measure because score updates mix rebuttal content with initial score position, paper-level consensus, reviewer confidence, and discussion dynamics. We study ICLR 2024-2025 using 73,000 reviewer trajectories with externally archived pre- and post-rebuttal scores, and use LLMs only as measurement instruments. Gemini Flash 3.0 predicts implied pre-rebuttal scores from score-stripped review text. The resulting text-score offset predicts later movement, with score-increase rates rising from 8.3% when text reads below the assigned score to 31.9% when it reads above. Claude Opus 4.6 induces, and outcome-blinded Gemini Flash 3.0 validates, a 44-feature taxonomy of resolved reviewer-author exchanges, where 23 features replicate across model and held-out year under Bonferroni correction. In the rebuttal-engaged benchmark (n=6,705), initial-review structure already predicts much score movement (AUC=0.747, minimal AUC=0.696), while adding the resolved exchange raises AUC to 0.804. Rebuttals can move scores, but measurable movement is bounded by initial-review structure, and robust exchange signals are mostly rebuttal failure modes.

cs.DL

Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods

Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extracted research question from each of 1,000 recent arXiv computer-science papers and compare the resulting methodology suggestions against a paper-derived experimental inventory. Since we provide only the research question, the differences we measure reflect initial suggestions and not how optimal those suggestions are. We extract structured method features from both sources, map them into a shared taxonomy, and quantify divergence across multiple taxonomy dimensions including model provider, dataset task type, and evaluation metric type. The strongest imbalance appears in provider choice, with Jensen-Shannon divergence about 3-5x larger than any other taxonomy dimension. Other/Academic single-occurrence models are underrepresented by 23-24 percentage points, while reused academic/community models are slightly overrepresented (4-6pp). LLMs also suggest a much narrower range of methods overall: the effective number of model entities contracts from 1,232 to 59-96, and inter-LLM rank correlations (0.55-0.68) generally exceed LLM-to-paper correlations (0.33-0.56), so the distortions are largely shared across models. Popularity baselines, BM25 retrieval calibration, and paper-level similarity tests confirm that the outputs are query-specific responses, but filtered through a narrower set of options. Researchers who rely on LLM suggestions without cross-checking therefore risk narrowing their methodological search space toward a more concentrated default.

cs.CL

On the Spectral Region of n-Cycle Stochastic Matrices

For every $n$, we determine the complete eigenvalue region of the $n$-cycle stochastic family. For $n\ge 2$, write $A_n(\alpha)$ for the matrix indexed by $\mathbb Z/n\mathbb Z$ with $$ (A_n(\alpha))_{j,j}=\alpha_j,\qquad (A_n(\alpha))_{j,j+1}=1-\alpha_j,\qquad 0\le \alpha_j<1, $$ and all other entries zero, and set $C_n=\{A_n(\alpha):\alpha\in[0,1)^n\}$. Writing $\Sigma_n$ for the corresponding spectral union, the trivial cases are $\Sigma_1=\{1\}$ and $\Sigma_2=[-1,1]$. For $n\ge 3$, we give an explicit description of $\Sigma_n$ in angular coordinates $m=\mathrm{Arg}(\lambda)$ and $M=\mathrm{Arg}(\lambda-1)$. Under the map $$ \Lambda(m,M)=\frac{\sin M}{\sin(M-m)}e^{im}, $$ the upper half of $\Sigma_n$ is the image of a finite union of $K=\lfloor(n-1)/2\rfloor$ vertical angular sectors. Its exposed boundary is an alternating chain of Jensen chords, arising from the Jensen-equality lines $M=\phi_k$, and algebraic one-loop arcs joining the relevant roots of unity to $0$; the lower boundary is obtained by complex conjugation. The real spectral part is $[-1,1]$ for even $n$ and $(0,1]$ for odd $n$. The proof is independent of Karpelevich's theorem and reduces the two-monomial characteristic equation to sharp argument bounds on a simplex, obtained by Jensen, majorization, and finite visibility arguments.

math.RA

Neutrino Fingerprints: Image-Based Encodings of IceCube Events for CNN Direction Reconstruction

Reconstructing the direction of incoming neutrinos in the IceCube Neutrino Observatory is an important problem in astrophysics. The public IceCube--Neutrinos in Deep Ice Kaggle competition provided 140 million simulated events to benchmark reconstruction techniques. To address this challenge from a novel perspective we introduce neutrino fingerprints compact $72 \times 72 \times 3$ images in which each pixel represents a single detector, with pulse timing and charge statistics encoded as color channels. This representation transforms sparse, irregular pulse data into dense images suitable for convolutional processing. Our ResNet18 model achieves a mean angular error of $1.10$ rad, indicating that convolutional networks trained on fingerprints rival more complex architectures while offering an effective, interpretable baseline for IceCube event reconstruction.

astro-ph.IM

One in Eight OpenAlex Abstracts Has Integrity Issues

Scientific abstracts are increasingly used as primary data in computational metascience research, yet the quality of these abstracts in widely used bibliographic databases has not been systematically examined. We assess the integrity of 10,000 randomly sampled English-language journal abstracts from OpenAlex using a two-stage annotation protocol combining human expert review and large language model classification. We identify seven distinct failure modes and find that 12\% of abstracts have integrity issues, with insufficient content and misplaced metadata being the most prevalent. We discuss implications for downstream research and describe a forthcoming community portal to support collective annotation efforts.

cs.DL

Inverse design of waveguide grating mode converters using artificial neural networks

Machine learning techniques, notably various deep neural network methods, are instrumental in processing extensive and intricate data sets in engineering and scientific fields. This paper shows how deep neural networks can inversely design cascaded-mode converting systems, particularly the waveguide gratings that implement selective mode conversion upon reflection. Neural networks can map the grating's physical features to scattering parameters of the modes reflected from the grating. The trained networks can then be utilized to inversely design the gratings based on the desired values of the scattering parameters. The process of the inverse design involves using the technique of gradient descent of a defined loss function. Minimizing this loss function leads to calculating more accurate features fulfilling the desired scattering parameters.

physics.optics

Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments

Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressing as generative AI undermines the reliability of take-home assessments, shifting emphasis toward supervised, in-class evaluation. We present a scalable, end-to-end workflow for LLM-assisted grading of short, pen-and-paper assessments. The workflow spans (1) constructing solution keys, (2) developing detailed rubric-style grading keys used to guide the LLM, and (3) a grading procedure that combines automated scanning and anonymisation, multi-pass LLM scoring, automated consistency checks, and mandatory human verification. We deploy the system in two undergraduate mathematics courses using six low-stakes in-class tests. Empirically, LLM assistance reduces grading time by approximately 23% while achieving agreement comparable to, and in several cases tighter than, fully manual grading. Occasional model errors occur but are effectively contained by the hybrid design. Overall, our results show that carefully embedded human-in-the-loop LLM grading can substantially reduce workload while maintaining fairness and accuracy.

cs.CY

Scalable Classification of Course Information Sheets Using Large Language Models: A Reusable Institutional Method for Academic Quality Assurance

Purpose: Higher education institutions face increasing pressure to audit course designs for generative AI (GenAI) integration. This paper presents an end-to-end method for using large language models (LLMs) to scan course information sheets at scale, identify where assessments may be vulnerable to student use of GenAI tools, validate system performance through iterative refinement, and operationalise results through direct stakeholder communication and effort. Method: We developed a four-phase pipeline: (0) manual pilot sampling, (1) iterative prompt engineering with multi-model comparison, (2) full production scan of 4,684 Bachelor and Master course information sheets (Academic Year 2024-2025) from the Vrije Universiteit Brussel (VUB) with automated report generation and email distribution to teaching teams (91.4% address-matched) using a three-tier risk taxonomy (Clear risk, Potential risk, Low risk), and (3) longitudinal re-scan of 4,675 sheets after the next catalogue release. Results: Five iterations of prompt refinement achieved 87% agreement with expert labels. GPT-4o was selected for production based on superior handling of ambiguous cases involving internships and practical components. The Year 1 scan classified 60.3% of courses as Clear risk, 15.2% as Potential risk, and 24.5% as Low risk. Year 2 comparison revealed substantial shifts in risk distributions, with improvements most pronounced in practice-oriented programmes. Implications: The method enables institutions to rapidly transform heterogeneous catalogue data into structured and actionable intelligence. The approach is transferable to other audit domains (sustainability, accessibility, pedagogical alignment) and provides a template for responsible LLM deployment in higher education governance.

cs.LG

Ergodicity in reinforcement learning

In reinforcement learning, we typically aim to optimize the expected value of the sum of rewards an agent collects over a trajectory. However, if the process generating these rewards is non-ergodic, the expected value, i.e., the average over infinitely many trajectories with a given policy, is uninformative for the average over a single, but infinitely long trajectory. Thus, if we care about how the individual agent performs during deployment, the expected value is not a good optimization objective. In this paper, we discuss the impact of non-ergodic reward processes on reinforcement learning agents through an instructive example, relate the notion of ergodic reward processes to more widely used notions of ergodic Markov chains, and present existing solutions that optimize long-term performance of individual trajectories under non-ergodic reward dynamics.

cs.LG

Probing Graph Neural Network Activation Patterns Through Graph Topology

Curvature notions on graphs provide a theoretical description of graph topology, highlighting bottlenecks and denser connected regions. Artifacts of the message passing paradigm in Graph Neural Networks, such as oversmoothing and oversquashing, have been attributed to these regions. However, it remains unclear how the topology of a graph interacts with the learned preferences of GNNs. Through Massive Activations, which correspond to extreme edge activation values in Graph Transformers, we probe this correspondence. Our findings on synthetic graphs and molecular benchmarks reveal that MAs do not preferentially concentrate on curvature extremes, despite their theoretical link to information flow. On the Long Range Graph Benchmark, we identify a systemic \textit{curvature shift}: global attention mechanisms exacerbate topological bottlenecks, drastically increasing the prevalence of negative curvature. Our work reframes curvature as a diagnostic probe for understanding when and why graph learning fails.

cs.LG

Space-time beams with tunable orbital group velocity toward plasma superradiance

Light springs are space-time beams that have a helical wavepacket. Due to this special property, light springs result into a rotating pulse when intercepting a plane lying orthogonal to their propagation direction. Associated to this, we introduce here the orbital group velocity, an additional tunable property of light springs. The orbital group velocity quantifies the speed of the light spring intensity rotation, distinctly from the conventional longitudinal group velocity, which describes the motion of the wavepacket envelope along its propagation axis. We demonstrate experimentally by tunable Fourier synthesis that the orbital group velocity can assume sub- and superluminal values, thus becoming a new platform for synthetic motion studies and control of laser-matter interactions. Particularly, in the superluminal regime, when interacting with a thin overdense plasma, we reveal by particle-in-cell simulations that the light spring unlocks superradiant radiation, due to the coherent excitation of the electrons in the plasma acting as a quasiparticle. This superradiant source inherits the ultrafast temporal dynamics of the light springs while emitting in the terahertz region, thus creating a new source of terahertz radiation controlled by the properties of spatiotemporal coupling of the laser. Therefore, spatiotemporal tuning of light springs is at the frontier of controlling laser-matter interaction and generating new tunable sources of radiation.

physics.optics

Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)

Large Language Models (LLMs) are increasingly used as scientific copilots, but evidence on their role in research-level mathematics remains limited, especially for workflows accessible to individual researchers. We present early evidence for vibe-proving with a consumer subscription LLM through an auditable case study that resolves Conjecture 20 of Ran and Teng (2024) on the exact nonreal spectral region of a 4-cycle row-stochastic nonnegative matrix family. We analyze seven shareable ChatGPT-5.2 (Thinking) threads and four versioned proof drafts, documenting an iterative pipeline of generate, referee, and repair. The model is most useful for high-level proof search, while human experts remain essential for correctness-critical closure. The final theorem provides necessary and sufficient region conditions and explicit boundary attainment constructions. Beyond the mathematical result, we contribute a process-level characterization of where LLM assistance materially helps and where verification bottlenecks persist, with implications for evaluation of AI-assisted research workflows and for designing human-in-the-loop theorem proving systems.

cs.AI

Probing the Trajectories of Reasoning Traces in Large Language Models

Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and decision commitment evolve along a reasoning trajectory, and whether intermediate trace segments provide answer-relevant information beyond generic length or stylistic effects. Here, we propose a protocol to systematically probe the trajectories of reasoning traces in LLMs by 1) generating a model's reasoning trace, 2) truncating it at fixed token-percentiles, and 3) injecting each partial trace back into the model (or a different model) to measure the induced distribution over answer choices via next-token probabilities. We apply this protocol to the open-source Qwen3-4B/-8B/-14B and gpt-oss-20b/-120b models across the multiple-choice GPQA Diamond and MMLU-Pro benchmarks. We find that accuracy and decision commitment consistently increase as the percentage of provided reasoning tokens grows. These gains are primarily driven by relevant content in the model generation rather than context length or generic "reasoning style" effects. Stronger models often backtrack successfully from incorrect partial traces, but immediate answers often remain anchored in the weaker model's incorrect response. More broadly, we show that trajectory probing provides diagnostics for efficient and safer deployment of reasoning models as the measurements can inform practical trace-handling and monitoring policies that improve reliability without assuming intermediate tokens are inherently faithful explanations.

cs.LG