SearcharxivSearch

arXiv subjects

Jialiang Huang

Publications and source records attributed to Jialiang Huang.

7 recordsLinked to original sources

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

cs.CL

Thin film synthesis of SrZn2P2 with SrI2 post-annealing for enhanced crystallinity and optoelectronic quality

Ternary Zintl phosphides are promising light-absorbing semiconductors for thin-film optoelectronic applications, but strategies for controlling their microstructure and optoelectronic quality remain underexplored. Here, we report the synthesis of phase-pure SrZn2P2 thin films using radio-frequency co-sputtering in a PH3 + Ar atmosphere and investigate the impact of post-growth processing on their structural and optical properties. Grazing-incidence X-ray scattering and Raman spectroscopy confirm the formation of crystalline SrZn2P2 films over a finite compositional window. Optical measurements reveal strong absorption near the direct-band-gap energy (~1.8 eV) and near-band-edge photoluminescence. Further, we have studied the effects of chemically compatible halide-assisted annealing. It is found that SrI2 treatments lead to pronounced grain growth and reduced diffraction peak broadening while preserving phase purity, in contrast to rapid thermal or forming-gas annealing. Notably, annealing with SrI2 at 450 °C significantly enhances both the intensity and spatial uniformity of the photoluminescence, thus connecting the observed microstructural consolidation with improved radiative recombination. Our study demonstrates that halide-assisted annealing provides an effective pathway for microstructural control in SrZn2P2 thin films and highlights a generalizable processing strategy for advancing Zintl phosphide semiconductors toward optoelectronic applications.

cond-mat.mtrl-sci

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.

cs.CL

TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes

Serverless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv-X, a co-designed integration of the serverless platform with the operating system and CXL/RDMA-based remote memory pools. TrEnv-X's core innovations are repurposable sandboxes, which can be shared across different functions to decrease the associated creation overhead, and OS-level memory templates, which enable rapid state restoration from CXL/RDMA-based remote memory pools. To further demonstrate TrEnv-X's versatility, we generalize its design from traditional containers for microVM-based agent workloads and introduce new optimizations, including browser sharing and a page cache bypassing mechanism. Our evaluation shows that TrEnv-X achieves up to 7x reduction in P99 latency and 48% memory savings for container-based functions. When applied to LLM agents, it reduces the P99 latency by up to 58% and memory usage by 61% compared to state-of-the-art systems like E2B.

cs.DC

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2) Scalable Reinforcement Learning Framework: By implementing a robust reinforcement learning protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par with Gemini-3.0-Pro, achieving gold-medal performance in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI). (3) Large-Scale Agentic Task Synthesis Pipeline: To integrate reasoning into tool-use scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This methodology facilitates scalable agentic post-training, yielding substantial improvements in generalization and instruction-following robustness within complex, interactive environments.

cs.CL

Photoswitchable exceptional points derived from bound states in the continuum

Bound states in the continuum (BICs) and exceptional points (EPs), as two distinct physical singularities represented by complex frequencies in non-Hermitian systems, have garnered significant attention and clear definitions in their respective fields in recent years. They share overlapping applications in areas such as high-sensitivity sensing and laser emission. However, the transition between the two, inspired by these intersections, remains largely unexplored. In this work, we reveal the transition process in a non-Hermitian two-mode system, evolving from one bound singularity to a two-dimensional exceptional ring, where the EP is the coalescent state of the quasi-Friedrich-Wintgen (FW)-BIC. This phenomenon is experimentally validated through pored dielectric metasurfaces in terahertz band. Furthermore, external pumping induced photocarriers as the dissipative perturbation, facilitates the breaking of degeneracy in the complex eigenfrequency and enables dynamic EP switching. Finally, we experimentally demonstrate a switchable terahertz beam deflection driven by the phase singularities of the EP. These findings are instrumental in advancing the development of compact devices for sensing and wavefront control within non-Hermitian systems.

physics.optics

Stylized Story Generation with Style-Guided Planning

Current storytelling systems focus more ongenerating stories with coherent plots regard-less of the narration style, which is impor-tant for controllable text generation. There-fore, we propose a new task, stylized story gen-eration, namely generating stories with speci-fied style given a leading context. To tacklethe problem, we propose a novel generationmodel that first plans the stylized keywordsand then generates the whole story with theguidance of the keywords. Besides, we pro-pose two automatic metrics to evaluate theconsistency between the generated story andthe specified style. Experiments demonstratesthat our model can controllably generateemo-tion-driven orevent-driven stories based onthe ROCStories dataset (Mostafazadeh et al.,2016). Our study presents insights for stylizedstory generation in further research.

cs.CL