SearcharxivSearch

arXiv subjects

Takumi Ito

Publications and source records attributed to Takumi Ito.

At least 19 recordsLinked to original sources

Foreground Characterization and Mitigation in the Observations of the CD/EoR with the SKA

The Square Kilometre Array (SKA), with its unprecedented sensitivity, frequency coverage, and large collecting area, is poised to revolutionize our understanding of the Cosmic Dawn (CD) and Epoch of Reionization (EoR) epochs marking the formation of the first luminous sources and the subsequent reionization of the intergalactic medium (IGM). However, detecting the faint redshifted 21-cm signal from neutral hydrogen remains one of the foremost challenges in observational cosmology, as it is buried beneath bright foregrounds from Galactic synchrotron radiation, free-free emission, and extragalactic point sources that are 4-5 orders of magnitude stronger than the cosmological signal. In this chapter, we highlight the key components and characteristics of these foregrounds and review ongoing efforts to model, characterize, and mitigate them. We emphasize how the SKA-Low AA* configuration, through its optimized array design, wide field of view, and improved calibration accuracy, enhances our capacity to suppress foreground contamination and recover the cosmological signal. The SKA Observatory Foreground Challenge plays a pivotal role in this effort by bringing together the global EoR/CD community to develop, compare, and validate foreground removal pipelines using realistic simulated datasets. Building on the experience of existing pathfinders such as LOFAR, MWA, and HERA, these collaborative initiatives are helping refine statistical and machine learning-based approaches for signal recovery. Together, these advancements are laying the groundwork for the SKA to probe the thermal and ionization history of the early Universe with unprecedented precision.

astro-ph.CO

Feasible Force Set Shaping for a Payload-Carrying Platform Consisting of Tiltable Multiple UAVs Connected Via Passive Hinge Joints

This paper presents a method for shaping the feasible force set of a payload-carrying platform composed of multiple Unmanned Aerial Vehicles (UAVs) and proposes a control law that leverages the advantages of this shaped force set. The UAVs are connected to the payload through passively rotatable hinge joints. The joint angles are controlled by the differential thrust produced by the rotors, while the total force generated by all the rotors is responsible for controlling the payload. The shape of the set of the total force depends on the tilt angles of the UAVs, which allows us to shape the feasible force set by adjusting these tilt angles. This paper aims to ensure that the feasible force set encompasses the required shape, enabling the platform to generate force redundantly -meaning in various directions. We then propose a control law that takes advantage of this redundancy.

cs.RO

Linear Representations of Hierarchical Concepts in Language Models

We investigate how and to what extent hierarchical relations (e.g., Japan $\subset$ Eastern Asia $\subset$ Asia) are encoded in the internal representations of language models. Building on Linear Relational Concepts, we train linear transformations specific to each hierarchical depth and semantic domain, and characterize representational differences associated with hierarchical relations by comparing these transformations. Going beyond prior work on the representational geometry of hierarchies in LMs, our analysis covers multi-token entities and cross-layer representations. Across multiple domains we learn such transformations and evaluate in-domain generalization to unseen data and cross-domain transfer. Experiments show that, within a domain, hierarchical relations can be linearly recovered from model representations. We then analyze how hierarchical information is encoded in representation space. We find that it is encoded in a relatively low-dimensional subspace and that this subspace tends to be domain-specific. Our main result is that hierarchy representation is highly similar across these domain-specific subspaces. Overall, we find that all models considered in our experiments encode concept hierarchies in the form of highly interpretable linear representations.

cs.CL

Decision Support System for Technology Opportunity Discovery: An Application of the Schwartz Theory of Basic Values

Discovering technology opportunities (TOD) remains a critical challenge for innovation management, especially in early-stage development where consumer needs are often unclear. Existing methods frequently fail to systematically incorporate end-user perspectives, resulting in a misalignment between technological potentials and market relevance. This study proposes a novel decision support framework that bridges this gap by linking technological feasibility with fundamental human values. The framework integrates two distinct lenses: the engineering-based Technology Readiness Levels (TRL) and Schwartz's theory of basic human values. By combining these, the approach enables a structured exploration of how emerging technologies may satisfy diverse user motivations. To illustrate the framework's feasibility and insight potential, we conducted exploratory workshops with general consumers and internal experts at Sony Computer Science Laboratories, Inc., analyzing four real-world technologies (two commercial successes and two failures). Two consistent patterns emerged: (1) internal experts identified a wider value landscape than consumers (vision gap), and (2) successful technologies exhibited a broader range of associated human values (value breadth), suggesting strategic foresight may underpin market success. This study contributes both a practical tool for early-stage R\&D decision-making and a theoretical link between value theory and innovation outcomes. While exploratory in scope, the findings highlight the promise of value-centric evaluation as a foundation for more human-centered technology opportunity discovery.

cs.HC

On Entity Identification in Language Models

We analyze the extent to which internal representations of language models (LMs) identify and distinguish mentions of named entities, focusing on the many-to-many correspondence between entities and their mentions. We first formulate two problems of entity mentions -- ambiguity and variability -- and propose a framework analogous to clustering quality metrics. Specifically, we quantify through cluster analysis of LM internal representations the extent to which mentions of the same entity cluster together and mentions of different entities remain separated. Our experiments examine five Transformer-based autoregressive models, showing that they effectively identify and distinguish entities with metrics analogous to precision and recall ranging from 0.66 to 0.9. Further analysis reveals that entity-related information is compactly represented in a low-dimensional linear subspace at early LM layers. Additionally, we clarify how the characteristics of entity representations influence word prediction performance. These findings are interpreted through the lens of isomorphism between LM representations and entity-centric knowledge structures in the real world, providing insights into how LMs internally organize and use entity information.

cs.CL

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Pre-training large language models (LLMs) faces significant memory challenges due to the large size of model parameters. We introduce STaged parameter-Efficient Pre-training (STEP), which integrates parameter-efficient tuning techniques with model growth. We conduct experiments on pre-training LLMs of various sizes and demonstrate that STEP achieves up to a 53.9% reduction in maximum memory requirements compared to vanilla pre-training while maintaining equivalent performance. Furthermore, we show that the model by STEP performs comparably to vanilla pre-trained models on downstream tasks after instruction tuning.

cs.CL

Energy-Aware Task Allocation for Teams of Multi-mode Robots

This work proposes a novel multi-robot task allocation framework for robots that can switch between multiple modes, e.g., flying, driving, or walking. We first provide a method to encode the multi-mode property of robots as a graph, where the mode of each robot is represented by a node. Next, we formulate a constrained optimization problem to decide both the task to be allocated to each robot as well as the mode in which the latter should execute the task. The robot modes are optimized based on the state of the robot and the environment, as well as the energy required to execute the allocated task. Moreover, the proposed framework is able to encompass kinematic and dynamic models of robots alike. Furthermore, we provide sufficient conditions for the convergence of task execution and allocation for both robot models.

cs.RO

Design and Control of a VTOL Aerial Vehicle Tilting its Rotors Only with Rotor Thrusts and a Passive Joint

This paper presents a novel VTOL UAV that owns a link connecting four rotors and a fuselage by a passive joint, allowing the control of the rotor's tilting angle by adjusting the rotors' thrust. This unique structure contributes to eliminating additional actuators, such as servo motors, to control the tilting angles of rotors, resulting in the UAV's weight lighter and simpler structure. We first derive the dynamical model of the newly designed UAV and analyze its controllability. Then, we design the controller that leverages the tiltable link with four rotors to accelerate the UAV while suppressing a deviation of the UAV's angle of attack from the desired value to restrain the change of the aerodynamic force. Finally, the validity of the proposed control strategy is evaluated in simulation study.

eess.SY

Reference-free Evaluation Metrics for Text Generation: A Survey

A number of automatic evaluation metrics have been proposed for natural language generation systems. The most common approach to automatic evaluation is the use of a reference-based metric that compares the model's output with gold-standard references written by humans. However, it is expensive to create such references, and for some tasks, such as response generation in dialogue, creating references is not a simple matter. Therefore, various reference-free metrics have been developed in recent years. In this survey, which intends to cover the full breadth of all NLG tasks, we investigate the most commonly used approaches, their application, and their other uses beyond evaluating models. The survey concludes by highlighting some promising directions for future research.

cs.CL

KAgoshima Galactic Object survey with Nobeyama 45-metre telescope by Mapping in Ammonia lines (KAGONMA): Discovery of parsec-scale CO depletion in the Canis Major star-forming region

In observational studies of infrared dark clouds, the number of detections of CO freeze-out onto dust grains (CO depletion) at pc-scale is extremely limited, and the conditions for its occurrence are, therefore, still unknown. We report a new object where pc-scale CO depletion is expected. As a part of Kagoshima Galactic Object survey with Nobeyama 45-m telescope by Mapping in Ammonia lines (KAGONMA), we have made mapping observations of NH3 inversion transition lines towards the star-forming region associated with the CMa OB1 including IRAS 07077-1026, IRAS 07081-1028, and PGCC G224.28-0.82. By comparing the spatial distributions of the NH3 (1,1) and C18O (J=1-0), an intensity anti-correlation was found in IRAS 07077-1026 and IRAS 07081-1028 on the ~1 pc scale. Furthermore, we obtained a lower abundance of C18O at least in IRAS 07077-1026 than in the other parts of the star-forming region. After examining high density gas dissipation, photodissociation, and CO depletion, we concluded that the intensity anti-correlation in IRAS 07077-1026 is due to CO depletion. On the other hand, in the vicinity of the centre of PGCC G224.28-0.82, the emission line intensities of both the NH3 (1,1) and C18O (J=1-0) were strongly detected, although the gas temperature and density were similar to IRAS 07077-1026. This indicates that there are situations where C18O (J=1-0) cannot trace dense gas on the pc scale and implies that the conditional differences that C18O (J=1-0) can and cannot trace dense gas are unclear.

astro-ph.GA

Sub-kpc scale gas density histogram of the Galactic molecular gas: a new statistical method to characterise galactic-scale gas structures

To understand physical properties of the interstellar medium (ISM) on various scales, we investigate it at parsec resolution on the kiloparsec scale. Here, we report on the sub-kpc scale Gas Density Histogram (GDH) of the Milky Way. The GDH is a density probability distribution function (PDF) of the gas volume density. Using this method, we are free from an identification of individual molecular clouds and their spatial structures. We use survey data of $^{12}$CO and $^{13}$CO ($J$=1-0) emission in the Galactic plane ($l = 10^{\circ}$-$50^{\circ}$) obtained as a part of the FOREST Unbiased Galactic plane Imaging survey with the Nobeyama 45m telescope (FUGIN). We make a GDH for every channel map of $2^{\circ} \times 2^{\circ}$ area, including the blank sky component, and without setting cloud boundaries. This is a different approach from previous works for molecular clouds. The GDH fits well to a single or double log-normal distribution, which we name the low-density log-normal (L-LN) and high-density log-normal (H-LN) components, respectively. The multi-log-normal components suggest that the L-LN and H-LN components originate from two different stages of structure formation in the ISM. Moreover, we find that both the volume ratios of H-LN components to total ($f_{\mathrm{H}}$) and the width of the L-LN along the gas density axis ($σ_{\rm{L}}$) show coherent structure in the Galactic-plane longitude-velocity diagram. It is possible that these GDH parameters are related to strong galactic shocks and other weak shocks in the Milky Way.

astro-ph.GA

Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP

We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/less reproducible. We present our results and findings, which include that just 13\% of papers had (i) sufficiently low barriers to reproduction, and (ii) enough obtainable information, to be considered for reproduction, and that all but one of the experiments we selected for reproduction was discovered to have flaws that made the meaningfulness of conducting a reproduction questionable. As a result, we had to change our coordinated study design from a reproduce approach to a standardise-then-reproduce-twice approach. Our overall (negative) finding that the great majority of human evaluations in NLP is not repeatable and/or not reproducible and/or too flawed to justify reproduction, paints a dire picture, but presents an opportunity for a rethink about how to design and report human evaluations in NLP.

cs.CL

Exploring the Robustness of Large Language Models for Solving Programming Problems

Using large language models (LLMs) for source code has recently gained attention. LLMs, such as Transformer-based models like Codex and ChatGPT, have been shown to be highly capable of solving a wide range of programming problems. However, the extent to which LLMs understand problem descriptions and generate programs accordingly or just retrieve source code from the most relevant problem in training data based on superficial cues has not been discovered yet. To explore this research question, we conduct experiments to understand the robustness of several popular LLMs, CodeGen and GPT-3.5 series models, capable of tackling code generation tasks in introductory programming problems. Our experimental results show that CodeGen and Codex are sensitive to the superficial modifications of problem descriptions and significantly impact code generation performance. Furthermore, we observe that Codex relies on variable names, as randomized variables decrease the solved rate significantly. However, the state-of-the-art (SOTA) models, such as InstructGPT and ChatGPT, show higher robustness to superficial modifications and have an outstanding capability for solving programming problems. This highlights the fact that slight modifications to the prompts given to the LLMs can greatly affect code generation performance, and careful formatting of prompts is essential for high-quality code generation, while the SOTA models are becoming more robust to perturbations.

cs.CL

Lower Perplexity is Not Always Human-Like

In computational psycholinguistics, various language models have been evaluated against human reading behavior (e.g., eye movement) to build human-like computational models. However, most previous efforts have focused almost exclusively on English, despite the recent trend towards linguistic universal within the general community. In order to fill the gap, this paper investigates whether the established results in computational psycholinguistics can be generalized across languages. Specifically, we re-examine an established generalization -- the lower perplexity a language model has, the more human-like the language model is -- in Japanese with typologically different structures from English. Our experiments demonstrate that this established generalization exhibits a surprising lack of universality; namely, lower perplexity is not always human-like. Moreover, this discrepancy between English and Japanese is further explored from the perspective of (non-)uniform information density. Overall, our results suggest that a cross-lingual evaluation will be necessary to construct human-like computational models.

cs.CL

Gate voltage dependence of noise distribution in radio-frequency reflectometry in gallium arsenide quantum dots

We investigate gate voltage dependence of electrical readout noise in high-speed rf reflectometry using gallium arsenide quantum dots. The fast Fourier transform spectrum from the real time measurement reflects build-in device noise and circuit noise including the resonator and the amplifier. We separate their noise spectral components by model analysis. Detail of gate voltage dependence of the flicker noise is investigated and compared to the charge sensor sensitivity. We point out that the dominant component of the readout noise changes by the measurement integration time.

cond-mat.mes-hall

Langsmith: An Interactive Academic Text Revision System

Despite the current diversity and inclusion initiatives in the academic community, researchers with a non-native command of English still face significant obstacles when writing papers in English. This paper presents the Langsmith editor, which assists inexperienced, non-native researchers to write English papers, especially in the natural language processing (NLP) field. Our system can suggest fluent, academic-style sentences to writers based on their rough, incomplete phrases or sentences. The system also encourages interaction between human writers and the computerized revision system. The experimental results demonstrated that Langsmith helps non-native English-speaker students write papers in English. The system is available at https://emnlp-demo.editor. langsmith.co.jp/.

cs.CL

Language Models as an Alternative Evaluator of Word Order Hypotheses: A Case Study in Japanese

We examine a methodology using neural language models (LMs) for analyzing the word order of language. This LM-based method has the potential to overcome the difficulties existing methods face, such as the propagation of preprocessor errors in count-based methods. In this study, we explore whether the LM-based method is valid for analyzing the word order. As a case study, this study focuses on Japanese due to its complex and flexible word order. To validate the LM-based method, we test (i) parallels between LMs and human word order preference, and (ii) consistency of the results obtained using the LM-based method with previous linguistic studies. Through our experiments, we tentatively conclude that LMs display sufficient word order knowledge for usage as an analysis tool. Finally, using the LM-based method, we demonstrate the relationship between the canonical word order and topicalization, which had yet to be analyzed by large-scale experiments.

cs.CL

Diamonds in the Rough: Generating Fluent Sentences from Early-Stage Drafts for Academic Writing Assistance

The writing process consists of several stages such as drafting, revising, editing, and proofreading. Studies on writing assistance, such as grammatical error correction (GEC), have mainly focused on sentence editing and proofreading, where surface-level issues such as typographical, spelling, or grammatical errors should be corrected. We broaden this focus to include the earlier revising stage, where sentences require adjustment to the information included or major rewriting and propose Sentence-level Revision (SentRev) as a new writing assistance task. Well-performing systems in this task can help inexperienced authors by producing fluent, complete sentences given their rough, incomplete drafts. We build a new freely available crowdsourced evaluation dataset consisting of incomplete sentences authored by non-native writers paired with their final versions extracted from published academic papers for developing and evaluating SentRev models. We also establish baseline performance on SentRev using our newly built evaluation dataset.

cs.CL