SearcharxivSearch

arXiv subjects

Haruto Sato

Publications and source records attributed to Haruto Sato.

3 recordsLinked to original sources

Behaviorally Effective LoRA Writes Are Sparse and Structured

Low-rank adaptation fixes the rank of the update, but it does not identify which parts of a trained write actually carry behavior. We study that question directly and show that behaviorally effective LoRA writes are sparse, structured, and far more concentrated than the raw low-rank parameterization suggests. We use Learned-Basis LoRA, a learned-basis continuation recipe, to expose that structure. The recipe warms up an unconstrained adapter, converts its learned write columns into a module-wise orthonormal basis, freezes that basis, and continues training inside the constrained parameterization. Across 14 exact switches from unconstrained to constrained form, held-out accuracy is unchanged at the conversion step and reconstructed write matrices differ by at most 0.25% relative Frobenius error. Same-state continuation then shows that the same trained checkpoint develops differently under different write subspaces, establishing write geometry as a causal state variable. A no-retraining projection test shows that useful write signal stays inside the learned write space and largely disappears from random or frozen-activation PCA controls. The concentration pattern is strong at both local and global scales. Across GSM8K, MathQA, and AQuA, per-module top-k continuation reaches its optimum at k in {2, 4} in all twelve seed-level cases we test. A stricter global ranking test shows that learned top-16 and top-32 subsets outperform matched random subsets, especially on GSM8K/Qwen and MathQA/Qwen. Single-direction ablations further reveal a sparse set of late q_proj, o_proj, and down_proj components with outsized behavioral impact.

cs.CL

Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA

Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported answer appear plausible. We study selective answering through evidence sufficiency boundaries: for the same question, a model should abstain under unsupported or partially supported context, answer when the context first becomes sufficient, and keep the answer stable when redundant evidence is added. We introduce Evidence Sufficiency Boundary Training, a generation-native training framework that constructs ordered evidence chains and supervises the abstain-to-answer transition directly. The method combines level supervision, a boundary flip margin, post-boundary stability, and answer recall protection. We build evidence chains from HotpotQA, 2WikiMultiHopQA, and MuSiQue, then evaluate models with chain metrics, raw QA utility, and unsupported-answer rates on external non-answerable sets. With Qwen2.5-3B-Instruct and LoRA adaptation, Evidence Sufficiency Boundary Training gives the strongest boundary localization among the tested systems, with flip accuracy of 0.807 compared with 0.781 for a token-level abstention baseline. It also achieves the lowest overall unsupported-answer rate on external non-answerable evaluation, 0.095 compared with 0.101 for the same baseline, while retaining competitive raw QA F1. The results show that grounded selective answering improves when training marks the evidence level where refusal should give way to answering.

cs.CL

Thermodynamic stability of twisted domains in AgCrSe$_{2}$ thin films grown on lattice-matched YSZ(111) substrate

Control of structural domains in epitaxial thin films of functional materials is a fundamental technique to utilize their intrinsic physical and chemical properties in solid-state devices. In this study, we report on suppression of twisted-domain formation in thin-film growth of polar magnetic semiconductor AgCrSe$_{2}$ using pulsed-laser deposition. In exploring concomitant optimized growth temperature and Ag/Cr composition ratio of supply, we find the critical growth temperature ($T\mathrm{_{sub}}$) for obtaining single 60$^{\circ}$ domain in c-axis oriented AgCrSe$_{2}$ thin film on a lattice-matched (111) plane of the yttria-stabilized zirconia substrate. At temperatures below and above the critical $T\mathrm{_{sub}}$, metastable 0$^{\circ}$ domain in addition to the 60$^{\circ}$ domain emerges, indicating delicate energy balance of thermodynamic stability for obtaining the single-domain structure. Surface structural analysis using time-of-flight low-energy atom scattering spectroscopy reveals the presence of two polar orientations along $+Z$ and $-Z$ directions. These findings provide valuable insights into the thin-film growth mechanisms for a family of two-dimensional compounds with rhombohedral lattices.

cond-mat.mtrl-sci