SearcharxivSearch

arXiv subjects

Takaaki Fukai

Publications and source records attributed to Takaaki Fukai.

6 recordsLinked to original sources

An Experimental Evaluation of Accurate Scheduling and Hardware Timestamping on NVIDIA ConnectX NICs

High-precision packet transmission is becoming increasingly important in deterministic networking applications, including 5G fronthaul and Time-Sensitive Networking (TSN). Recent NVIDIA ConnectX network interface cards (NICs) provide Accurate Scheduling, a hardware-assisted mechanism for transmitting Ethernet frames at designated times, as part of their 5T for 5G feature set. They also provide hardware timestamping for received and transmitted frames. Although these functions are expected to satisfy the stringent timing requirements of 5G fronthaul, little public information is available regarding their timing accuracy and performance characteristics. This paper presents an experimental characterization of the Accurate Scheduling and hardware timestamping capabilities of the NVIDIA ConnectX-7 NIC. Using an FPGA-based measurement platform with deterministic frame generation and nanosecond-resolution timestamping, we first evaluate the precision of the receive and transmit hardware timestamps and then evaluate the transmission timing accuracy of Accurate Scheduling. The experimental results show that the receive and transmit hardware timestamps exhibit a measured variation of approximately $\pm$7-8 ns when compared across independently clocked Ethernet entities. Furthermore, Accurate Scheduling transmits approximately 99% of frames within $\pm$900 ns of the specified transmission time, while occasional outliers of up to approximately 5 us are observed. These results indicate that Accurate Scheduling is well suited for applications with latency requirements on the order of several tens of microseconds, such as 5G fronthaul, whereas its timing accuracy is insufficient for highly deterministic TSN applications, which typically require transmission timing accuracy on the order of several to several tens of nanoseconds.

cs.NI

VolTune: A Fine-Grained Runtime Voltage Control Architecture for FPGA Systems

The rapid emergence of edge computing platforms and large-scale data centers has made power efficiency a primary design constraint, particularly for data-intensive and AI-driven workloads. Field-programmable gate arrays (FPGAs) are increasingly adopted due to their flexibility and potential for energy-efficient acceleration. However, FPGA supply voltages are typically fixed at design time using conservative margins, limiting the ability to adapt power consumption to runtime conditions. This paper presents VolTune, an open-source runtime voltage control architecture that enables runtime tuning of FPGA supply voltages through FPGA-integrated control logic that abstracts low-level PMBus operations. VolTune provides both hardware-based and software-based control paths, allowing designers to balance deterministic low-latency operation against programmability. In the presented prototype, the hardware-based control path achieves a measured end-to-end voltage transition latency of 2.3 ms, while the controller adds under 2% static power overhead and under 2% FPGA resource overhead. As a representative case study, VolTune is evaluated on the GTX transceiver supply rail of a Kintex-7 platform. The results show that runtime voltage tuning exposes a bounded operating region with clear trade-offs between energy efficiency and reliability, and achieves up to approximately 29.3% rail-power reduction at 10.0 Gbps when allowing BER up to 10e-6. These results show that FPGA-integrated runtime voltage control can provide practical energy savings with low integration overhead.

cs.AR

NecoFuzz: Effective Fuzzing of Nested Virtualization via Fuzz-Harness Virtual Machines

Nested virtualization is now widely supported by major cloud vendors, allowing users to leverage virtualization-based technologies in the cloud. However, supporting nested virtualization significantly increases host hypervisor complexity and introduces a new attack surface in cloud platforms. While many prior studies have explored hypervisor fuzzing, none has explicitly addressed nested virtualization due to the challenge of generating effective virtual machine (VM) instances with a vast state space as fuzzing inputs. We present NecoFuzz, the first fuzzing framework that systematically targets nested virtualization-specific logic in hypervisors. NecoFuzz synthesizes executable fuzz-harness VMs with internal states near the boundary between valid and invalid, guided by an approximate model of hardware-assisted virtualization specifications. Since vulnerabilities in nested virtualization often stem from incorrect handling of unexpected VM states, this specification-guided, boundary-oriented generation significantly improves coverage of security-critical code across different hypervisors. We implemented NecoFuzz on Intel VT-x and AMD-V by extending AFL++ to support fuzz-harness VMs. NecoFuzz achieved 84.7% and 74.2% code coverage for nested virtualization-specific code on Intel VT-x and AMD-V, respectively, and uncovered six previously unknown vulnerabilities across three hypervisors, including two assigned CVEs.

cs.OS

METICULOUS: An FPGA-based Main Memory Emulator for System Software Studies

Due to the scaling problem of the DRAM technology, non-volatile memory devices, which are based on different principle of operation than DRAM, are now being intensively developed to expand the main memory of computers. Disaggregated memory is also drawing attention as an emerging technology to scale up the main memory. Although system software studies need to discuss management mechanisms for the new main memory designs incorporating such emerging memory systems, there are no feasible memory emulation mechanisms that efficiently work for large-scale, privileged programs such as operating systems and hypervisors. In this paper, we propose an FPGA-based main memory emulator for system software studies on new main memory systems. It can emulate the main memory incorporating multiple memory regions with different performance characteristics. For the address region of each memory device, it emulates the latencies, bandwidths and bit-flip error rates of read/write operations, respectively. The emulator is implemented at the hardware module of an off-the-self FPGA System-on-Chip board. Any privileged/unprivileged software programs running on its powerful 64-bit CPU cores can access emulated main memory devices at a practical speed through the exactly same interface as normal DRAM main memory. We confirmed that the emulator transparently worked for CPU cores and successfully changed the performance of a memory region according to given emulation parameters; for example, the latencies measured by CPU cores were exactly proportional to the latencies inserted by the emulator, involving the minimum overhead of approximately 240 ns. As a preliminary use case, we confirmed that the emulator allows us to change the bandwidth limit and the inserted latency individually for unmodified software programs, making discussions on latency sensitivity much easier.

cs.AR

Analyzing I/O Performance of a Hierarchical HPC Storage System for Distributed Deep Learning

Today, deep learning is an essential technology for our life. To solve more complex problems with deep learning, both sizes of training datasets and neural networks are increasing. To train a model with large datasets and networks, distributed deep neural network (DDNN) training technique is necessary. For large-scale DDNN training, HPC clusters are a promising computation environment. In large-scale DDNN on HPC clusters, I/O performance is critical because it is becoming a bottleneck. Most flagship-class HPC clusters have hierarchical storage systems. For designing future HPC storage systems, it is necessary to quantify the performance improvement effect of the hierarchical storage system on the workloads. This paper demonstrates the quantitative performance analysis of the hierarchical storage system for DDNN workload in a flagship-class supercomputer. Our analysis shows how much performance improvement and volume increment of the storage will be required to meet the performance goal.

cs.DC

MLPerf HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems

Scientific communities are increasingly adopting machine learning and deep learning models in their applications to accelerate scientific insights. High performance computing systems are pushing the frontiers of performance with a rich diversity of hardware resources and massive scale-out capabilities. There is a critical need to understand fair and effective benchmarking of machine learning applications that are representative of real-world scientific use cases. MLPerf is a community-driven standard to benchmark machine learning workloads, focusing on end-to-end performance metrics. In this paper, we introduce MLPerf HPC, a benchmark suite of large-scale scientific machine learning training applications driven by the MLCommons Association. We present the results from the first submission round, including a diverse set of some of the world's largest HPC systems. We develop a systematic framework for their joint analysis and compare them in terms of data staging, algorithmic convergence, and compute performance. As a result, we gain a quantitative understanding of optimizations on different subsystems such as staging and on-node loading of data, compute-unit utilization, and communication scheduling, enabling overall $>10 \times$ (end-to-end) performance improvements through system scaling. Notably, our analysis shows a scale-dependent interplay between the dataset size, a system's memory hierarchy, and training convergence that underlines the importance of near-compute storage. To overcome the data-parallel scalability challenge at large batch sizes, we discuss specific learning techniques and hybrid data-and-model parallelism that are effective on large systems. We conclude by characterizing each benchmark with respect to low-level memory, I/O, and network behavior to parameterize extended roofline performance models in future rounds.

cs.LG