Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 

README.md

Overview of Japanese LLMs

[ English | Français | 日本語 ]

📖 Please visit the more readable Web version

The content of this README is available in a more readable format at llm-jp.github.io/awesome-japanese-llm/en. We recommend viewing the Web version to avoid table display issues and layout problems.

A list of publicly available LLMs trained with a focus on Japanese, along with their evaluation benchmarks, maintained by volunteers from various sources like academic papers and other public resources.

::: warning Caution

  1. We can't guarantee the accuracy or completeness of any information here.
  2. Some information is based on conjecture and might not reflect your specific use case.
  3. While many models are released under permissive licenses like MIT or Apache 2.0, some are subject to more restrictive terms including non-commercial use clauses (e.g CC BY-NC-SA 4.0) or other stipulations. :::

Please point out any errors on the issues page. Feel free to contribute directly with a pull request.

::: details Table of Contents {open} [[toc]] :::

Text Generation Models

For multimodal models, see below.

Models built from scratch

General purpose

Release Year Architecture Max Context Length Training Data Developer License / Terms of Use
Sarashina2-8x70B 2024 MoE
(8x70b (465b))
8,192 Sparse Upcycling on Sarashina2 (70B) SB Intuitions Sarashina Model NonCommercial License
LLM-jp-3 172B 2024 Llama
(172b, 172b-instruct2, 172b-instruct3)
4,096 Pre-training: llm-jp-corpus-v3
(2.1T tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, magpie-sft-v1.0, Daring-Anteater, FLAN, ichikara-instruction-format, AutoMultiTurnByCalm3-22B, ramdom-to-fixed-multiturn-Calm3, wizardlm8x22b-logical-math-coding-sft-ja, wizardlm8x22b-logical-math-coding-sft_additional-ja, Synthetic-JP-EN-Coding-Dataset-567k
DPO (instruct3 only): aya-ja-evol-inst, ac-self-inst
Research and Development Center for Large Language Models Pre-trained model: LLM-jp-3 172B Terms of Use
Post-trained model: llm-jp-3-172b-instruct3 Terms of Use
LLM-jp-3 172B beta2 2024 Llama
(172b-beta2, 172b-beta2-instruct2)
4,096 Pre-training: part of llm-jp-corpus-v3
(1.4T tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, magpie-sft-v1.0, Daring-Anteater, FLAN, ichikara-instruction-format, AutoMultiTurnByCalm3-22B, ramdom-to-fixed-multiturn-Calm3, wizardlm8x22b-logical-math-coding-sft-ja, wizardlm8x22b-logical-math-coding-sft_additional-ja, Synthetic-JP-EN-Coding-Dataset-567k
Research and Development Center for Large Language Models LLM-jp-3 172B beta2 Terms of Use
LLM-jp-3 172B beta1 2024 Llama
(172b-beta1, 172b-beta1-instruct)
4,096 Pre-training: part of llm-jp-corpus-v3
(0.7T tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, Dolly Dataset, OASST1, OASST2, Aya Dataset, ichikara-instruction-format, Daring-Anteater, FLAN
Research and Development Center for Large Language Models LLM-jp-3 172B beta1 Terms of Use
LLM-jp-3 172B alpha 2024 Llama
(172b-alpha1, 172b-alpha1-instruct, 172b-alpha2, 172b-alpha2-instruct)
4,096 Pre-training: part of llm-jp-corpus-v3
(alpha1: 0.7T tokens, alpha2: 1.4T tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, Dolly Dataset, OASST1, OASST2, Aya Dataset, ichikara-instruction-format, Daring-Anteater, FLAN
Research and Development Center for Large Language Models Apache 2.0
Stockmark-2-100B-Instruct-beta 2025 Llama
(100B-Instruct-beta, 100B-Instruct-beta-AWQ)
4,096 Pre-training: 1.5T tokens
Instruction Tuning
DPO
Stockmark MIT
Stockmark-100b 2024 Llama
(100b, 100b-instruct-v0.1)
4,096 Pre-training: RedPajama, Japanese Wikipedia, Japanese mC4, Japanese CommonCrawl, Japanese Patent, Stockmark Web Corpus
(910B tokens)
Instruction Tuning (LoRA): ichikara-instruction
Stockmark MIT
PLaMo-100B-Pretrained 2024 Llama1
(100b)
4,096 Pre-training: Japanese CommonCrawl, RefinedWeb, undisclosed
(2.0T tokens)
Preferred Elements (Preferred Networks) PLaMo Non-Commercial License
LLM-jp-3.1 2025 Llama/MoE
(8x13b (73b), 8x13b (73b)-instruct4, 13b, 13b-instruct4, 1.8b, 1.8b-instruct4)
4,096 Pre-training: llm-jp-corpus-v3
(2.5T tokens)
Continual pre-training: instruction-response pairs
(90B tokens)
SFT + DPO
Research and Development Center for Large Language Models Apache 2.0
LLM-jp-3 MoE 2025 MoE
(8x1.8b (9.3b), 8x1.8b (9.3b)-instruct2, 8x1.8b (9.3b)-instruct3, 8x13b (73b), 8x13b (73b)-instruct2, 8x13b (73b)-instruct3)
4,096 Drop-Upcycling on LLM-jp-3 (1.8b, 13b) Research and Development Center for Large Language Models Apache 2.0
Sarashina2 2024 Llama
(7b, 13b, 70b)
7b, 13b: 4,096
70b: 8,192
Pre-training: Japanese Common Crawl, SlimPajama, StarCoder
(2.1T tokens)
SB Intuitions MIT
Sarashina1 2024 GPT-NeoX
(7b, 13b, 65b)
2,048 Pre-training: Japanese Common Crawl
(1T tokens)
SB Intuitions MIT
Tanuki-8×8B 2024 MoE (47b)
(v1.0, v1.0-AWQ, v1.0-GPTQ-4bit, v1.0-GPTQ-8bit, v1.0-GGUF)
4,096 Pre-training: various Web & synthetic datasets(1.7T tokens)
SFT, DPO: various synthetic datasets 2
Matsuo Lab LLM Development Project Apache 2.0
LLM-jp-4 32B-A3B 2026 Qwen3 MoE
(32b-a3b-base, 32b-a3b-thinking, 32b-a3b-thinking-gguf)
65,536 Pre-training + Mid-training: llm-jp-corpus-v4.1, llm-jp-corpus-midtraining-v2
(11.7T tokens)
SFT: llm-jp-4-thinking-sft-data
DPO: llm-jp-4-32b-a3b-thinking-dpo-data
Research and Development Center for Large Language Models Apache 2.0
llm-jp-4-kappa 2026 Qwen3 MoE
(32b-a3b-v0.1)
65,536 Post-trained on LLM-jp-4 32B-A3B (thinking) as follows
SFT: Nemotron-Cascade-2-SFT-Data, OpenMathReasoning, OpenCodeReasoning, Nemotron-SFT-Competitive-Programming-v2, llm-jp-4-thinking-sft-data, etc. (443,747 samples in total)
GRPO (math): 10,999 problems selected from AceReason-Math
GRPO (code): approx. 3.1k samples selected from SYNTHETIC-2
Third Intelligence Apache 2.03
PLaMo 3 2025 Gemma-based architecture
(2b-base, 8b-base, 31b-base)
4,096 Pre-training: English, Japanese, Code, Multilingual
(2b: 200B tokens, 8b: 1T tokens, 31b: 3T tokens)
Preferred Networks PLaMo community license
Way-PLaMo-3-8b-chat 2025 PLaMo 3-based (8b-chat) 4,096 Instruction Following SFT: Alpaca (51.7K), Dolly-15k-ja (15K) Individual (WayBob) PLaMo community license
CyberAgentLM3 (CALM3) 2024 Llama
(22b-chat, 22b-chat-selfimprove-experimental)
16,384 undisclosed
(2.0T tokens)
CyberAgent Apache 2.0
LLM-jp-3 13B instruct3 2025 Llama
(150m, 150m-instruct2, 150m-instruct3, 440m, 440m-instruct2, 440m-instruct3, 980m, 980m-instruct2, 980m-instruct3, 1.8b-instrcut2, 1.8b-instruct3, 3.7b-instruct2, 3.7b-instruct3, 7.2b-instruct2, 7.2b-instruct3, 13b-instruct2, 13b-instruct3)
4,096 Pre-training: llm-jp-corpus-v3
(2.1T tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, magpie-sft-v1.0, Daring-Anteater, FLAN, ichikara-instruction-format, AutoMultiTurnByCalm3-22B, ramdom-to-fixed-multiturn-Calm3, wizardlm8x22b-logical-math-coding-sft-ja, Synthetic-JP-EN-Coding-Dataset-567k
DPO (instruct3 only): aya-ja-evol-inst, ac-self-inst
Research and Development Center for Large Language Models Apache 2.0
LLM-jp-3 13B 2024 Llama
(1.8b, 1.8b-instruct, 3.7b, 3.7b-instruct, 7.2b, 7.2b-instruct, 13b, 13b-instruct)
4,096 Pre-training: llm-jp-corpus-v3
(2.1T tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, FLAN, ichikara-instruction-format, AutoMultiTurnByCalm3-22B, ramdom-to-fixed-multiturn-Calm3, wizardlm8x22b-logical-math-coding-sft_additional-ja, Synthetic-JP-EN-Coding-Dataset-567k
Research and Development Center for Large Language Models Apache 2.0
llm-jp-3-3.7b-instruct-EZO 2024 Llama
(3.7b-instruct-EZO-Common, 3.7b-instruct-EZO-Humanities)
4,096 additionally trained on LLM-jp-3 (3.7B) Axcxept Apache 2.0
LLM-jp-13B v2.0 2024 Llama
(13b-v2.0, 13b-instruct-full-dolly-ichikara_004_001_single-oasst-oasst2-v2.0, 13b-instruct-full-ac_001-dolly-ichikara_004_001_single-oasst-oasst2-v2.0, 13b-instruct-full-ac_001_16x-dolly-ichikara_004_001_single-oasst-oasst2-v2.0)
4,096 Pre-training: llm-jp-corpus-v2
(260B tokens)
Instruction Tuning: ichikara-instruction, AnswerCarefully Dataset, Dolly Dataset, OASST1, OASST2
LLM-jp Apache 2.0
Fugaku-LLM 2024 GPT
(13B, 13B-instruct, 13B-instruct-gguf)
2,048 Pre-training: undisclosed dataset
Instruction Tuning: OASST1, Dolly Dataset, GSM8K
Titech, Tohoku Univ., Fujitsu, RIKEN, Nagoya Univ., CyberAgent, Kotoba Technologies Fugaku-LLM Terms of Use
LLM-jp-13B v1.1 2024 GPT
(13b-instruct-lora-dolly_en-dolly_ja-ichikara_003_001-oasst_en-oasst_ja-v1.1, 13b-instruct-full-dolly_en-dolly_ja-ichikara_003_001-oasst_en-oasst_ja-v1.1, 13b-dpo-lora-hh_rlhf_ja-v1.1)
2,048 Instruction Tuning (LoRA or Full-parameter FT): Dolly Dataset, OASST1, ichikara-instruction
DPO (LoRA): HH RLHF
LLM-jp Apache 2.0
LLM-jp-13B 2023 GPT
(1.3b-v1.0, 13b-v1.0, 13b-instruct-full-jaster-v1.0, 13b-instruct-full-jaster-dolly-oasst-v1.0, 13b-instruct-full-dolly-oasst-v1.0, 13b-instruct-lora-jaster-v1.0, 13b-instruct-lora-jaster-dolly-oasst-v1.0, 13b-instruct-lora-dolly-oasst-v1.0)
2,048 Pre-training: llm-jp-corpus (Wikipedia, Japanese mC4, The Pile, Stack) (300B tokens)
Instruction Tuning (Full-parameter FT or LoRA): jaster, Dolly Dataset, OASST1
LLM-jp Apache 2.0
PLaMo-13B 2023 Llama4
(13b, 13b-instruct, 13b-instruct-nc)
base: 4,096
instruct, instruct-nc: 8,192
Pre-training: C4, Project Gutenberg, RedPajama, Japanese Wikipedia, Japanese mC4
(1.5T tokens)
Instruction Tuning: Dolly, HH RLHF, OASST1, wikinews (+Alpaca in NC model)
Preferred Networks Apache 2.0
(CC BY-NC 4.0 as for NC model)
Stockmark-13b 2023 Llama
(13b, 13b-instruct)
2,048 Pre-training: Japanese Wikipedia, Japanese CC-100, Japanese mC4, Japanese CommonCrawl, Japanese Patent, Stockmark Web Corpus
(220B tokens)
Instruction Tuning (LoRA): ichikara-instruction
Stockmark base: MIT
instruct: CC BY-NC-SA 4.0
Weblab-10B 2023 GPT-NeoX
(10b, 10b-instruction-sft)
2,048 Japanese mC4, The Pile
(600B tokens)
Instruction Tuning: Alpaca, FLAN
University of Tokyo Matsuo-Iwasawa Lab CC BY‑NC 4.0
LLM-jp-4 8B 2026 Llama
(8b-base, 8b-instruct, 8b-thinking, 8b-thinking-gguf)
65,536 Pre-training + Mid-training: llm-jp-corpus-v4.1, llm-jp-corpus-midtraining-v2
(11.7T tokens)
SFT: llm-jp-4-thinking-sft-data
DPO (thinking only): llm-jp-4-8b-thinking-dpo-data
Research and Development Center for Large Language Models Apache 2.0
PLaMo 2.1 8B 2025 hybrid architecture like Samba
(8b-cpt)
32,768 Training details undisclosed Preferred Networks PLaMo community license
PLaMo 2 8B 2025 hybrid architecture like Samba
(8b)
mainly Japanese and English data
(6T tokens)
Preferred Networks PLaMo community license
Tanuki-8B 2024 Tanuki (8b)
(v1.0, v1.0-AWQ, v1.0-GPTQ-4bit, v1.0-GPTQ-8bit, v1.0-GGUF)
4,096 Pre-training: various Web & synthetic datasets(1.3T tokens)
SFT, DPO: various synthetic datasets 2
Matsuo Lab LLM Development Project Apache 2.0
Japanese StableLM Alpha 2023 GPT-NeoX
(base-alpha-7b, instruct-alpha-7b, instruct-alpha-7b-v2)
2,048 Wikipedia, Japanese CC‑100, Japanese mC4, Japanese OSCAR, RedPajama, private datasets5
(750B tokens)
Instruction Tuning: Dolly, HH‑RLHF, wikinews, Alpaca (discarded in v2)
Stability AI base: Apache 2.0
instruct (v1): Research license
instruct (v2): Apache 2.0
CyberAgentLM2 (CALM2) 2023 Llama
(7b, 7b-chat, 7b-chat-dpo-experimental)
base: 4,096
chat: 32,768
publicly available Japanese and English datasets (details unknown)
(1.3T tokens)
DPO: Chatbot Arena Conversations JA (calm2) Dataset
CyberAgent Apache 2.0
(CC BY 4.0 as for DPO model)
OpenCALM 2023 GPT-NeoX
(small, medium, large, 1b(1.4b), 3b(2.7b), 7b(6.8b))
2,048 Japanese Wikipedia, Japanese mC4, Japanese CC‑100 CyberAgent CC BY‑SA 4.0
Stormy 2023 GPT-NeoX
(7b(6.8b))
2,048 OpenCALM fine-tuned on
llm-japanese-dataset v0 non-translation tasks
University of Tokyo Izumi Lab CC BY‑SA 4.0
ByGPT-JP 2025 Llama-based
(multi-lm-head-6.5b-alpha)
5,760 Subset of llm-jp-corpus-v3 (ja_cc, ja_warp_html, ja_warp_pdf, ja_wiki, kaken) Tohoku University NLP Group Apache 2.0
rinna GPT
(En-Ja Bilingual)
2023 GPT-NeoX
(4b(3.8b), 4b(3.8b)-8k, 4b(3.8b)-instruction-sft, 4b(3.8b)-instruction-ppo)
8k model: 8,192
others: 2,048
Wikipedia, Japanese CC‑100, Japanese C4, RedPajama, The Pile
(524B tokens)
Instruction Tuning: HH‑RLHF, FLAN
PPO: HH‑RLHF for reinforcement learning
8k: trained with long context
rinna MIT
japanese-large-lm 2023 GPT-NeoX
(1.7b, 3.6b, 1.7b-instruction-sft, 3.6b-instruction-sft)
2,048 Japanese Wikipedia, Japanese CC‑100, Japanese C4, Japanese OSCAR and private datasets
(650GB)
Instruction Tuning: OASST1
LINE Apache 2.0
rinna GPT
(Japanese only)
2023 GPT / GPT-NeoX
(xsmall, small, medium, 1b, neox-small, neox-3.6b-instruction-sft-v2, neox-3.6b-instruction-ppo)
≤ 2,048 Japanese Wikipedia, Japanese CC‑100
(1b and later models add
Japanese mC4)
Instruction Tuning: HH‑RLHF, FLAN, SHP
PPO: HH‑RLHF for reinforcement learning
rinna MIT
Sarashina2.2 2025 Llama
(0.5b, 0.5b-instruct-v0.1, 1b, 1b-instruct-v0.1, 3b, 3b-instruct-v0.1)
8,192 SB Intuitions MIT
RetrievaT5 2023 T5
(small (short), small (medium), small (long), base (short), base (medium), base (long), large (short), large (medium), large (long), xl(3b))
Japanese Wikipedia, Japanese mC4 Retrieva CC BY‑SA 4.0
Spiral-RetNet-3b-base 2024 RetNet
(3b)
2,048 Wikipedia, Japanese CC-100, CulturaX Spiral.AI MIT
kotomamba-2.8B 2024 Mamba
(2.8B-v1.0)
2,048 Japanese Wikipedia, Swallow Corpus, SlimPajama Kotoba Technologies Apache 2.0
ABEJA GPT 2022 GPT / GPT-NeoX
(large, neox-2.7b)
Japanese Wikipedia, Japanese CC‑100, Japanese OSCAR ABEJA MIT
PLaMo 2.1 2B 2025 Causal decoder-only transformer
(2b-cpt)
32,768 Training details undisclosed Preferred Networks PLaMo community license
Rakuten AI 2.0 mini 2025 Mistral
(mini(1.5b), mini(1.5b)-instruct)
131,072 Rakuten Apache 2.0
WasedaGPT 2022 GPT
(small, xl(1.5b))
Japanese Wikipedia, Japanese CC‑100 Waseda Kawahara Lab CC BY‑SA 4.0
StockmarkGPT 2023 GPT-NeoX
(1.4b)
Japanese Wikipedia (0.88B tokens), Japanese CC‑100 (10.5B tokens), private data (8.6B tokens) Stockmark MIT
YellowbackGPT 2021 GPT-NeoX
(1.3b)
Japanese Wikipedia, Japanese CC‑100, Japanese OSCAR Yellowback Apache 2.0
PLaMo 2 1B 2025 hybrid architecture like Samba
(1b)
mainly Japanese and English data
(4T tokens)
Preferred Elements (Preferred Networks) Apache 2.0
Sarashina2.1-1B 2024 Llama
(1b)
8,192 Japanese and English data on the web (10T tokens) SB Intuitions Sarashina Model NonCommercial License
colorfulscoop GPT 2021 GPT
(small)
Japanese Wikipedia Colorful Scoop CC BY‑SA 3.0
TitechGPT 2023 GPT
(medium, medium-reversed) 6
Japanese Wikipedia, Japanese CC‑100 Titech Okazaki Lab CC BY‑SA 4.0
KyotoUniversityGPT 2022 GPT
(small, medium, large)
Japanese Wikipedia (3.2GB), Japanese CC‑100 (85GB), Japanese OSCAR (54GB) Kyoto University Language Media Processing Lab CC BY‑SA 4.0
JapaneseBART 2023 BART
(base, large)
Japanese Wikipedia (18M sentences) Kyoto University Language Media Processing Lab CC BY‑SA 4.0
Megagon Labs T5 2021 T5
(base)
Japanese mC4 (782 GB), Japanese wiki40b (2 GB) Megagon Labs
(Recruit Co.,Ltd.)
Apache 2.0

Domain Specific

Release Year Domain Architecture Training Data Developer License
SIP-med-LLM/SIP-jmed-llm-3-8x13b-OP-32k-R0.1 2026 Medical MoE Continual pre-training on a medical corpus (78.3B tokens) added to LLM-jp-3 MoE (8x13b), context length extended to 32k, followed by Instruction Tuning Strategic Innovation Promotion Program (SIP) Phase 3 Project "Generative AI Utilization in the Construction of Integrated Healthcare Systems" Theme 1 "Development and Social Implementation of Open Medical LLM with Safety and Reliability" Research Group Apache 2.07
SIP-med-LLM/SIP-jmed-llm-3-8x13b-AC-32k-instruct 2025 Medical MoE Continual pre-training on a medical corpus (78.3B tokens) added to LLM-jp-3 MoE (8x13b), context length extended to 32k, followed by Instruction Tuning using PMC-Patients and others Strategic Innovation Promotion Program (SIP) Phase 3 Project "Generative AI Utilization in the Construction of Integrated Healthcare Systems" Theme 1 "Development and Social Implementation of Open Medical LLM with Safety and Reliability" Research Group SIP-jmed-llm-3-8x13b-AC-32k-instruct Terms of Use7
SIP-med-LLM/SIP-jmed-llm-3-8x13b-OP-4k-base 2025 Medical MoE Continual pre-training on a medical corpus (78.3B tokens) added to LLM-jp-3 MoE (8x13b) Strategic Innovation Promotion Program (SIP) Phase 3 Project "Generative AI Utilization in the Construction of Integrated Healthcare Systems" Theme 1 "Development and Social Implementation of Open Medical LLM with Safety and Reliability" Research Group Apache 2.07
SIP-med-LLM/SIP-jmed-llm-2-8x13b-OP-instruct 2025 Medical MoE Pre-trained on a medical corpus (44.2B tokens) added to LLM-jp-3 MoE (8x13b), followed by Instruction Tuning Strategic Innovation Promotion Program (SIP) Phase 3 Project "Generative AI Utilization in the Construction of Integrated Healthcare Systems" Theme 1 "Development and Social Implementation of Open Medical LLM with Safety and Reliability" Research Group Apache 2.07
SIP-med-LLM/SIP-jmed-llm-3-13b-OP-32k-R0.1 2026 Medical Llama Continual pre-training on a medical corpus (78.3B tokens) added to LLM-jp-3.1 (13b), context length extended to 32k, followed by Instruction Tuning Strategic Innovation Promotion Program (SIP) Phase 3 Project "Generative AI Utilization in the Construction of Integrated Healthcare Systems" Theme 1 "Development and Social Implementation of Open Medical LLM with Safety and Reliability" Research Group Apache 2.07
SIP-med-LLM/SIP-jmed-llm-3-13b-OP-4k-base 2025 Medical Llama Continual pre-training on a medical corpus (78.3B tokens) added to LLM-jp-3.1 (13b) Strategic Innovation Promotion Program (SIP) Phase 3 Project "Generative AI Utilization in the Construction of Integrated Healthcare Systems" Theme 1 "Development and Social Implementation of Open Medical LLM with Safety and Reliability" Research Group Apache 2.07
AscleLM-1-10B 2026 Medical Mamba-2 Hybrid
(based on Nemotron-H)
Pre-training: Nemotron-CC v1 (approx. 6.3T tokens), open-source Japanese/English/Chinese corpora, code, math, and reasoning data (approx. 8.5T tokens in total), and synthetic data derived from medical papers, medical textbooks, clinical guidelines, and exam questions University of Tokyo Matsuo-Iwasawa Lab AscleLM-1-10B Terms of Use8
Japanese Dialog Transformer 2021 Dialog Transformer Twitter Japanese reply pairs NTT Evaluation Licence
Japanese News BART 2023 Business BART (base) Japanese business news articles (21M articles) Stockmark MIT
AcademicBART 2023 Science BART (base) CiNii Japanese Papers Ehime University AI Lab Apache 2.0

Models built off non-Japanese LLMs (w/ continual pre-training on Japanese)

*Includes models that underwent post-training (SFT, DPO, RL, etc.) after continual pre-training.

General purpose

Release Year Base Model Training Data Developer License / Terms of Use
GPT-OSS Swallow 120B
(120B-SFT-v0.1, 120B-RL-v0.1, 120B-RL-v0.1-MXFP4)
2026 GPT-OSS (120b) Pre-training: Wikipedia, Swallow Corpus v3.2, Nemotron-CC, Cosmopedia, Laboro ParaCorpus, Swallow Math v2, Swallow Code v2
(419.4B tokens)
SFT: GPT-OSS-LMSYS-Chat-1M-Synth, Swallow-Nemotron-Post-Training-Dataset-v1
RL: allenai/Dolci-Think-RL-7B (Math subset)
Swallow Project Apache 2.0
Llama 3.3 Swallow 70B
(70B-v0.4, 70B-Instruct-v0.4)
2025 Llama 3.3 (70b) Pre-training: Wikipedia, DCLM-baseline-1.0, Swallow Corpus Version 2, Cosmopedia, Laboro ParaCorpus, FineMath-4+, Swallow Code Version 0.3
Instruction Tuning: Gemma-2-LMSYS-Chat-1M-Synth, Swallow-Magpie-Ultra-v0.1, Swallow-Gemma-Magpie-v0.1, Swallow-Code-v0.3-Instruct-style
Swallow Project Llama 3.3 Community License & Gemma Terms of Use
Llama 3.1 Swallow 70B
(70B-v0.1, 70B-Instruct-v0.1, 70B-Instruct-v0.3)
2024 Llama 3.1 (70b) Pre-training: The Stack v2, Wikipedia, DCLM-baseline-1.0, Swallow Corpus Version 2, Cosmopedia, Laboro ParaCorpus
Instruction Tuning: lmsys-chat-1m-synth-ja-wo-pii-and-template-instructions, lmsys-chat-1m-synth-en-wo-pii-and-template-instructions, filtered-magpie-ultra-ja, filtered-magpie-ultra-en, gemma-magpie
Swallow Project Llama 3.1 Community License
(Gemma Terms of Use is also applied to the Instruct model)
cyberagent/Llama-3.1-70B-Japanese-Instruct-2407 2024 Llama 3.1 (70b) undisclosed CyberAgent Llama 3.1 Community License
Llama 3 Swallow 70B
(70B-v0.1, 70B-Instruct-v0.1)
2024 Llama 3 (70b) Pre-training: Algebraic Stack, Wikipedia, RefinedWeb, Swallow Corpus, Cosmopedia, Laboro ParaCorpus, OpenWebMath
Instruction Tuning: OASST1 9
Swallow Project Llama 3 Community License
turing-motors/Llama-3-heron-brain-70B-v0.3 2024 Llama 3 (70b) additionally trained on Llama 3 Swallow 70B (details undisclosed) Turing Llama 3 Community License
Llama 3 Youko 70B
(70b, 70b-instruct, 70b-gptq, 70b-instruct-gptq)
2024 Llama 3 (70b) Pre-training: Wikipedia, Japanese C4, Japanese CC-100, Japanese OSCAR, The Pile, undisclosed dataset
(5B tokens)
Instruction Tuning: undisclosed datasetト10
rinna Llama 3 Community License
Swallow 70B
(70b-hf, 70b-instruct-hf, 70b-instruct-v0.1, 70b-NVE-hf, 70b-NVE-instruct-hf)
2023 Llama 2 (70b) Pre-training: Japanese Wikipedia, RefinedWeb, Swallow Corpus, The Pile
Instruction Tuning: Dolly Dataset, HH RLHF, OASST1
*v0.1: OASST1, OASST2
Swallow Project Llama 2 Community License
KARAKURI LM
(70b-v0.1, 70b-chat-v0.1)
2024 Llama 2 (70b) Pre-training: mC4, CC100, OSCAR, RedPajama, undisclosed dataset
(16B tokens)
SteerLM: OASST2, undisclosed dataset
KARAKURI Llama 2 Community License11
Japanese Stable LM Beta 70B
(base-beta-70b, instruct-beta-70b)
2023 Llama 2 (70b) Pre-training: Wikipedia, Japanese mC4, Japanese CC-100, Japanese OSCAR, SlimPajama(excluding Books3)
(100B tokens)
Instruction Tuning: Dolly Dataset, HH RLHF, OASST1
Stability AI Llama 2 Community License
Fujitsu-LLM-KG
(8x7B_cpt, 8x7B_inst-infer_v1, 8x7B_inst-infer_v2, 8x7B_inst-gen_ja, 8x7B_inst-gen_en)
2025 Mixtral-8x7B-Instruct-v0.1 (46.7b) Pre-training: Knowledge graph parallel corpus (synthesized from Shinra Project, Wikipedia, etc.) 2.1B tokens, total ~300B tokens
Instruction Tuning: Knowledge graph reasoning and generation task datasets
Fujitsu Apache 2.0
Swallow-MX 8x7B
(8x7b-NVE-v0.1)
2024 Mixtral-8x7B-Instruct-v0.1 (46.7b) Pre-training: Algebraic Stack, Japanese Wikipedia, RefinedWeb, Swallow Corpus, The Pile, The Vault Swallow Project Apache 2.0
KARAKURI LM 8x7B Instruct v0.1
(8x7b-instruct-v0.1)
2024 Mixtral-8x7B-Instruct-v0.1 (46.7b) trained Swallow-MX 8x7B on the following datasets: Dolly Dataset, OASST2, HelpSteer, glaive-code-assistant-v3, glaive-function-calling-v2, synthetic_text_to_sql, MetaMathQA, orca-math-word-problems-200k, rag-dataset-12000, rag-hallucination-dataset-1000, undisclosed dataset KARAKURI Apache 2.0 (?)12
KARAKURI LM 8x7B Chat v0.1
(8x7b-chat-v0.1)
2024 Mixtral-8x7B-Instruct-v0.1 (46.7b) trained Swallow-MX 8x7B on OASST2, HelpSteer, and undisclosed datasets using SteerLM KARAKURI Apache 2.0
ABEJA-Mixtral-8x7B-japanese
(8x7B-v0.1-japanese, 8x7B-Instruct-v0.1-japanese, 8x7B-Instruct-v0.1-japanese-alpha, 8x7B-Instruct-v0.1-japanese-alpha-merged)
2024 Mixtral-8x7B-Instruct-v0.1 (46.7b)
*The model without "Instruct" in its name is based on Mixtral-8x7B-v0.1
Pre-training: Japanese CC, Redpajama, undisclosed dataset
450B tokens)
ABEJA Apache 2.0
Qwen3 Swallow 32B
(32B-CPT-v0.2, 32B-SFT-v0.2, 32B-RL-v0.2, 32B-RL-v0.2-AWQ-INT4)
2026 Qwen3 (32b) Pre-training: Wikipedia, Swallow Corpus v3.2, Nemotron-CC, Cosmopedia, Laboro ParaCorpus, Swallow Math v2, Swallow Code v2
(209.7B tokens)
SFT: GPT-OSS-LMSYS-Chat-1M-Synth, Swallow-Nemotron-Post-Training-Dataset-v1
RL: allenai/Dolci-Think-RL-7B (Math subset)
Swallow Project Apache 2.0
ELYZA-Thinking-1.0-Qwen-32B
(32B)
2025 Qwen 2.5 (32b) Pre-training + SFT (Reasoning) ELYZA Apache 2.0
ELYZA-Shortcut-1.0-Qwen-32B
(32B)
2025 Qwen 2.5 (32b) Pre-training + SFT ELYZA Apache 2.0
ABEJA-Qwen2.5-32b-Japanese-v1.0
(v1.0)
2025 Qwen2.5-32B-Instruct (32b) Continual pre-training + SFT + DPO: ~20,000 synthetic and human-annotated datasets (specialized for extraction and reasoning) ABEJA Apache 2.0
Qwen2.5 Bakeneko 32B
(qwen2.5-bakeneko-32b, qwen2.5-bakeneko-32b-instruct, deepseek-r1-distill-qwen2.5-bakeneko-32b, qwq-bakeneko-32b)
2025 Qwen 2.5 (32b) rinna Apache 2.0
ABEJA-QwQ32b-Reasoning-Japanese-v1.0
(v1.0)
2025 Qwen 2.5 (32b) ABEJA-Qwen2.5-32b-Japanese-v0.1 + Chat Vector (from QwQ 32b) + continual pre-training ABEJA Apache 2.0
ABEJA-Qwen2.5-32b-Japanese-v0.1
(32b-Japanese-v0.1)
2025 Qwen 2.5 (32b) Pre-training: Common Crawl, Cosmopedia, undisclosed dataset
100B tokens)
+ Chat Vector
ABEJA Apache 2.0
neoAI-JP-QwQ-32B
(32B)
2025 Qwen 2.5 (32b) Continual pre-training: ~4B tokens from llm-jp-corpus v3
+ Chat Vector (from QwQ-32B)
neoAI Apache 2.0
neoAI-JP-DeepSeek-Qwen-32B
(32B)
2025 Qwen 2.5 (32b) Continual pre-training: ~4B tokens from llm-jp-corpus v3
+ Chat Vector (from DeepSeek-R1-Distill-Qwen-32B)
neoAI Apache 2.0
Qwen3 Swallow 30B-A3B
(30B-A3B-CPT-v0.2, 30B-A3B-SFT-v0.2, 30B-A3B-RL-v0.2, 30B-A3B-RL-v0.2-AWQ-INT4)
2026 Qwen3 (30b-A3B) Pre-training: Wikipedia, Swallow Corpus v3.2, Nemotron-CC, Cosmopedia, Laboro ParaCorpus, Swallow Math v2, Swallow Code v2
(209.7B tokens)
SFT: GPT-OSS-LMSYS-Chat-1M-Synth, Swallow-Nemotron-Post-Training-Dataset-v1
RL: allenai/Dolci-Think-RL-7B (Math subset)
Swallow Project Apache 2.0
Gemma-2-Llama Swallow 27B
(27b-pt-v0.1, 27b-it-v0.1)
2025 Gemma 2 (27b) Pre-training: Wikipedia, DCLM-baseline-1.0, Swallow Corpus Version 2, Cosmopedia, Laboro ParaCorpus, FineMath-4+, Swallow Code Version 0.3
Instruction Tuning: Gemma-2-LMSYS-Chat-1M-Synth, Swallow-Magpie-Ultra-v0.1, Swallow-Gemma-Magpie-v0.1
Swallow Project Llama 3.3 Community License & Gemma Terms of Use
GPT-OSS Swallow 20B
(20B-SFT-v0.1, 20B-RL-v0.1, 20B-RL-v0.1-MXFP4)
2026 GPT-OSS (20b) Pre-training: Wikipedia, Swallow Corpus v3.2, Nemotron-CC, Cosmopedia, Laboro ParaCorpus, Swallow Math v2, Swallow Code v2
(419.4B tokens)
SFT: GPT-OSS-LMSYS-Chat-1M-Synth, Swallow-Nemotron-Post-Training-Dataset-v1
RL: allenai/Dolci-Think-RL-7B (Math subset)
Swallow Project Apache 2.0
ABEJA-Qwen3-14B-Agentic-256k-v0.1
(v0.1)
2026 Qwen3 (14b) Continual pre-training: long-context synthetic data (extended context length from 128k to 256k via YaRN)
+ Reinforcement learning for enhanced agentic capabilities
ABEJA Apache 2.0
Nekomata 14B
(14b, 14b-instruction, 14b-gguf, 14b-instruction-gguf)
2023 Qwen (14b) Pre-training: Wikipedia, Japanese C4, Japanese CC-100, Japanese OSCAR, The Pile, undisclosed dataset
(66B tokens)
Instruction Tuning: Dolly Dataset, FLAN, subsets of llm-japanese-dataset
rinna Tongyi Qianwen LICENSE
Swallow 13B
(13b-hf, 13b-instruct-hf, 13b-instruct-v0.1, 13b-NVE-hf)
2023 Llama 2 (13b) Pre-training: Japanese Wikipedia, RefinedWeb, Swallow Corpus, The Pile
Instruction Tuning: Dolly Dataset, HH RLHF, OASST1
*v0.1: OASST1, OASST2
Swallow Project Llama 2 Community License
LEIA-Swallow-13B
(13b)
2024 Llama 2 (13b) additionally trained Swallow 13B using LEIA Individual (Ikuya Yamada, Ryokan Ri) Llama 2 Community License
ELYZA-japanese-Llama-2-13b
(13b, 13b-instruct, 13b-fast, 13b-fast-instruct)
2023 Llama 2 (13b) Pre-training: Japanese Wikipedia, Japanese OSCAR, and other crawled data
(18B tokens)
Instruction Tuning: undisclosed dataset
ELYZA Llama 2 Community License
cyberagent/Mistral-Nemo-Japanese-Instruct-2408 2024 Mistral NeMo (12b) undisclosed CyberAgent Apache 2.0
NVIDIA-Nemotron-Nano-9B-v2-Japanese
(9B)
2026 Nemotron-Nano (9b) Continual pre-training: Wikipedia, fineweb-2 Japanese, aozorabunko, sip3-ja-general-web-corpus, Nemotron-CC-v2.1, Nemotron-Pretraining-Specialized-v1
SFT: Tool Calling data seeded from Nemotron-Personas-Japan, Nemotron-Post-Training-v3
NVIDIA NVIDIA Nemotron Open Model License Agreement
Gemma-2-Llama Swallow 9B
(9b-pt-v0.1, 9b-it-v0.1)
2025 Gemma 2 (9b) Pre-training: Wikipedia, DCLM-baseline-1.0, Swallow Corpus Version 2, Cosmopedia, Laboro ParaCorpus, FineMath-4+, Swallow Code Version 0.3
Instruction Tuning: Gemma-2-LMSYS-Chat-1M-Synth, Swallow-Magpie-Ultra-v0.1, Swallow-Gemma-Magpie-v0.1
Swallow Project Llama 3.3 Community License & Gemma Terms of Use
Qwen3 Swallow 8B
(8B-CPT-v0.2, 8B-SFT-v0.2, 8B-RL-v0.2, 8B-RL-v0.2-AWQ-INT4)
2026 Qwen3 (8b) Pre-training: Wikipedia, Swallow Corpus v3.2, Nemotron-CC, Cosmopedia, Laboro ParaCorpus, Swallow Math v2, Swallow Code v2
(209.7B tokens)
SFT: GPT-OSS-LMSYS-Chat-1M-Synth, Swallow-Nemotron-Post-Training-Dataset-v1
RL: allenai/Dolci-Think-RL-7B (Math subset)
Swallow Project Apache 2.0
Llama 3.1 Swallow 8B
(8B-v0.1, 8B-Instruct-v0.1, 8B-v0.2, 8B-Instruct-v0.2, 8B-Instruct-v0.3, 8B-Instruct-v0.5)
2025 Llama 3.1 (8b) Pre-training: The Stack v2, Wikipedia, DCLM-baseline-1.0, Swallow Corpus Version 2, Cosmopedia, Laboro ParaCorpus
Instruction Tuning: lmsys-chat-1m-synth-ja-wo-pii-and-template-instructions, lmsys-chat-1m-synth-en-wo-pii-and-template-instructions, filtered-magpie-ultra-ja, filtered-magpie-ultra-en, gemma-magpie, Gemma-3-LMSYS-Chat-1M-Synth
Swallow Project Llama 3.1 Community License
(Gemma Terms of Use is also applied to the Instruct model)
Llama 3 Swallow 8B
(8B-v0.1, 8B-Instruct-v0.1)
2023 Llama 3 (8b) Pre-training: Algebraic Stack, Wikipedia, RefinedWeb, Swallow Corpus, Cosmopedia, Laboro ParaCorpus, OpenWebMath
Instruction Tuning: OASST1 9
Swallow Project Llama 3 Community License
turing-motors/Llama-3-heron-brain-8B-v0.3 2024 Llama 3 (8b) additionally trained on Llama 3 Swallow 8B (details undisclosed) Turing Llama 3 Community License
Llama 3 Youko 8B
(8b, 8b-instruct, 8b-gptq, 8b-instruct-gptq)
2024 Llama 3 (8b) Pre-training: Wikipedia, Japanese C4, Japanese CC-100, Japanese OSCAR, The Pile, undisclosed dataset
(22B tokens)
Instruction Tuning10: Aya Dataset (Japanese subset), FLAN, Dolly Dataset, HH RLHF, OASST1, OASST2, MetaMathQA, CodeAlpaca Dataset, undisclosed dataset
DPO: HelpSteer, HelpSteer2, undisclosed dataset
rinna Llama 3 Community License
Llama 3 ELYZA JP 8B
(8B, 8B-GGUF, 8B-AWQ)
2024 Llama 3 (8b) undisclosed ELYZA Llama 3 Community License
Llama 3 neoAI 8B Chat v0.1
(8B-Chat-v0.1)
2024 Llama 3 (8b) undisclosed neoAI Llama 3 Community License
Llama 3 tedllm
(v0)
2024 Llama 3 (8b) Pre-training: Japanese generic corpus Tokyo Electron Device Llama 3 Community License
ELYZA-Shortcut-1.0-Qwen-7B
(7B)
2025 Qwen 2.5 (7b) Pre-training + SFT ELYZA Apache 2.0
ELYZA-Diffusion-1.0-Dream-7B
(Base-7B, Instruct-7B)
2026 Dream (7b) Pre-training: Japanese text (62B tokens)
Instruction Tuning: Japanese instruction data (
0.18B tokens)
ELYZA Apache 2.0
Swallow 7B
(7b-hf, 7b-instruct-hf, 7b-instruct-v0.1, 7b-NVE-hf, 7b-NVE-instruct-hf, 7b-plus-hf)
2023 Llama 2 (7b) Pre-training: Japanese Wikipedia, RefinedWeb, Swallow Corpus, The Pile
Instruction Tuning: Dolly Dataset, HH RLHF, OASST1
*v0.1: OASST1, OASST2
Swallow Project Llama 2 Community License
LEIA-Swallow-7B
(7b)
2024 Llama 2 (7b) additionally trained Swallow 7B using LEIA Individual (Ikuya Yamada, Ryokan Ri) Llama 2 Community License
ELYZA-japanese-Llama-2-7b
(7b, 7b-instruct, 7b-fast, 7b-fast-instruct)
2023 Llama 2 (7b) Pre-training: Japanese Wikipedia, Japanese OSCAR, and other crawled data
(18B tokens)
Instruction Tuning: undisclosed dataset
ELYZA Llama 2 Community License
Youri 7B
(7b, 7b-instruction, 7b-chat, 7b-gptq, 7b-instruction-gptq, 7b-chat-gptq)
2023 Llama 2 (7b) Pre-training: Wikipedia, Japanese C4, Japanese CC-100, Japanese OSCAR, The Pile, undisclosed dataset
(40B tokens)
Instruction Tuning: Dolly Dataset, FLAN, subsets of llm-japanese-dataset
rinna Llama 2 Community License
houou-7b
(instruction-7b-v1, instruction-7b-v2, instruction-7b-v3)
2023 Llama 2 (7b) Instruction-tuned Youri 7B (base) on ichikara-instruction MoneyForward Llama 2 Community License
Japanese Stable LM Beta 7B
(base-beta-7b, base-ja_vocab-beta-7b, instruct-beta-7b, instruct-ja_vocab-beta-7b)
2023 Llama 2 (7b) Pre-training: Wikipedia, Japanese mC4, Japanese CC-100, Japanese OSCAR, SlimPajama(excluding Books3)
(100B tokens)
Instruction Tuning: Dolly Dataset, HH RLHF, OASST1
Stability AI Llama 2 Community License
SambaLingo-Japanese
(Base, Chat)
2024 Llama 2 (7b) Pre-training: CulturaX
Instruction Tuning: ultrachat_200k
DPO: ultrafeedback, cai-conversation-harmless
SambaNova Systems Llama 2 Community License (?)12
blue-lizard
(blue-lizard)
2024 Llama 2 (7b) undisclosed Deepreneur Llama 2 Community License
Swallow-MS 7B
(7b-v0.1, 7b-instruct-v0.1)
2024 Mistral-7B-v0.1 (7b) Pre-training: Algebraic Stack, Japanese Wikipedia, RefinedWeb, Swallow Corpus, The Pile
Instruction Tuning: Dolly Dataset, OASST1
Swallow Project Apache 2.0
Rakuten AI 2.0
(8x7B, 8x7B-instruct)
2025 Mistral-7B-v0.1 (7b) Rakuten Apache 2.0
RakutenAI-7B
(7B, 7B-instruct, 7B-chat)
2024 Mistral-7B-v0.1 (7b) Pre-training: undisclosed
Instruction Tuning: Dolly Dataset, OASST1, datasets converted from the train split of NLU datasets (like jaster), undisclosed dataset
Rakuten Apache 2.0
Japanese Stable LM Gamma 7B
(base-gamma-7b, instruct-gamma-7b)
2023 Mistral-7B-v0.1 (7b) Pre-training: Wikipedia, Japanese mC4, Japanese CC-100, Japanese OSCAR, SlimPajama(excluding Books3)
(100B tokens)
Instruction Tuning: Dolly Dataset, HH RLHF, wikinews subset of llm-japanese-dataset
Stability AI Apache 2.0
ChatNTQ JA 7B
(7b-v1.0)
2024 Mistral-7B-v0.1 (7b) Instruction-tuned Japanese Stable LM Gamma 7B (base) on their own datasets NTQ Solution Apache 2.0
Shisa Gamma 7B
(7b-v1)
2023 Mistral-7B-v0.1 (7b) Instruction-tuned Japanese Stable LM Gamma 7B (base) on ultra-orca-boros-en-ja AUGMXNT Apache 2.0 (?)12
Shisa 7B
(base-7b-v1, 7b-v1)
2023 Mistral-7B-v0.1 (7b) Pre-training: shisa-pretrain-en-ja-v1 (8B tokens)
Instruction Tuning & DPO: ultra-orca-boros-en-ja, shisa-en-ja-dpo-v1
AUGMXNT Apache 2.0 (?)12
Karasu
(7B, 7B-chat, 7B-chat-plus, 7B-chat-plus-unleashed)
2024 Mistral-7B-v0.1 (7b) Additionally trained Shisa 7B (base) on Aozora Bunko, Japanese Law Precedent Dataset, Japanese Wikipedia, Japanese domain webscrapes from the Japanese subset of CulturaX, UltraChat 200k
(7B tokens)
Instruction Tuning: ultra-orca-boros-en-ja-v1, OASST1, ShareGPT, undisclosed dataset
Lightblue Apache 2.0 (?)12
Nekomata 7B
(7b, 7b-instruction, 7b-gguf, 7b-instruction-gguf)
2023 Qwen (7b) Pre-training: Wikipedia, Japanese C4, Japanese CC-100, Japanese OSCAR, The Pile, undisclosed dataset
(66B tokens)
Instruction Tuning: Dolly Dataset, FLAN, subsets of llm-japanese-dataset
rinna Tongyi Qianwen LICENSE
lightblue/japanese-mpt-7b 2023 MPT (7b) Japanese mC4 Lightblue Apache 2.0
Japanese Stable LM 3B-4E1T
(3b-4e1t-base, 3b-4e1t-instruct)
2024 StableLM-3B-4E1T (3b) Pre-training: Wikipedia, Japanese mC4, Japanese CC-100, Japanese OSCAR, SlimPajama(excluding Books3)
(100B tokens)
Instruction Tuning: Dolly Dataset, HH RLHF, wikinews subset of llm-japanese-dataset
Stability AI Apache 2.0
kotomamba-2.8B-CL 2024 mamba-2.8b-slimpj
(2.8b)
Japanese Wikipedia, Swallow Corpus, SlimPajama Kotoba Technologies Apache 2.0
Gemma-2-Llama Swallow 2B
(2b-pt-v0.1, 2b-it-v0.1)
2025 Gemma 2 (2b) Pre-training: Wikipedia, DCLM-baseline-1.0, Swallow Corpus Version 2, Cosmopedia, Laboro ParaCorpus, FineMath-4+, Swallow Code Version 0.3
Instruction Tuning: Gemma-2-LMSYS-Chat-1M-Synth, Swallow-Magpie-Ultra-v0.1, Swallow-Gemma-Magpie-v0.1
Swallow Project Llama 3.3 Community License & Gemma Terms of Use
Gemma 2 Baku 2B
(2b, 2b-it)
2024 Gemma 2 (2b) Pre-training: Wikipedia, Japanese C4, Japanese CC-100, Japanese OSCAR, The Pile, undisclosed dataset
(80B tokens)
OPRO: undisclosed dataset 13
rinna Gemma Terms of Use
Japanese Stable LM 2 1.6B
(base, instruct)
2024 Stable LM 2 1.6B (1.6b) Pre-training: Wikipedia, CulturaX
Instruction Tuning: jaster, ichikara-instruction, alpaca-gpt4-japanese, ultra-orca-boros-en-ja-v1
Stability AI STABILITY AI NON-COMMERCIAL RESEARCH COMMUNITY LICENSE
TinySwallow-1.5B
(1.5B, 1.5B-Instruct, 1.5B-Instruct-q4f32_1-MLC, 1.5B-Insturct-GGUF)
2025 Qwen2.5 (1.5b) Pre-training: trained using the TAID method (with Qwen2.5 (32b) as the teacher model)
Instruction Tuning: Gemma-2-LMSYS-Chat-1M-Synth, swallow-magpie-ultra-v0.1, swallow-gemma-magpie-v0.1
Sakana AI, Swallow Project Apache 2.0
EQUES/OpenRS3-GRPO-ja 2025 Qwen2.5 (1.5b) GRPO training on TinySwallow-1.5B-Instruct with kunishou/OpenMathInstruct-1-1.8m-ja EQUES Inc. ?
EQUES/TinyDeepSeek-JP-1.5B 2025 Qwen2.5 (1.5b) TAID distillation on TinySwallow-1.5B-Instruct with EQUES/japanese_ultrachat_6.6k EQUES Inc. Apache 2.0
EQUES/TinySwallow-Stratos-1.5B 2025 Qwen2.5 (1.5b) Reasoning enhancement on TinySwallow-1.5B-Instruct with Bespoke-Stratos-35k EQUES Inc. Apache 2.0
karasu-1.1B 2023 TinyLlama (1.1b) Pre-training: Japanese OSCAR, Japanese mC4
(3B tokens)
Lightblue Apache 2.0

Domain specific

Release Year Domain Base Model Training Data Developer License
Weblab-MedLLM-GLM-4.7 2026 Medicine GLM-4.7 (355b-A32B) Continual pre-training on proprietary synthetic data derived from medical papers, medical textbooks, clinical guidelines, and exam questions, followed by post-training University of Tokyo Matsuo-Iwasawa Lab MIT14
Weblab-MedLLM-Qwen3-235B
(Instruct, Thinking)
2026 Medicine Qwen3 (235b-A22B) Continual pre-training on proprietary synthetic data derived from medical papers, medical textbooks, clinical guidelines, and exam questions, followed by post-training University of Tokyo Matsuo-Iwasawa Lab Apache 2.014
Weblab-MedLLM-gpt-oss-120b 2026 Medicine GPT-OSS (120b) Continual pre-training on proprietary synthetic data derived from medical papers, medical textbooks, clinical guidelines, and exam questions, followed by post-training University of Tokyo Matsuo-Iwasawa Lab Apache 2.014
Medical-GPT-OSS-Swallow-120B 2026 Medicine GPT-OSS (120b) Continual pre-training of GPT-OSS Swallow 120B (RL) on a mixture emphasizing medical-domain text (biomedical literature, medical synthetic data, medical QA, clinical guideline-style text) Swallow Project Apache 2.07
pfnet/Preferred-MedLLM-Qwen-72B 2025 Medicine Qwen2.5 (72b) Continual pre-training on a proprietary corpus of medical-related text Preferred Networks Qwen LICENSE
Llama3-Preferred-MedSwallow-70B
(70B)
2024 Medicine Llama 3 (70b) Trained on a proprietary PFN medical dataset (e.g., explanations from pre-2017 Japanese national medical licensing exams) Preferred Networks Llama 3 Community License
AIgroup-CVM-utokyohospital/MedSwallow-70b 2024 Medicine Llama 2 (70b) Instruction tuning on a Japanese-translated USMLE dataset University of Tokyo Hospital Department of Cardiovascular Medicine AI Group CC BY-NC-SA 4.0
Medical-Qwen3-Swallow-32B 2026 Medicine Qwen3 (32b) Continual pre-training of Qwen3 Swallow 32B (RL) on a mixture emphasizing medical-domain text (biomedical literature, medical synthetic data, medical QA, clinical guideline-style text) Swallow Project Apache 2.07
Medical-Qwen3-Swallow-30B-A3B 2026 Medicine Qwen3 (30b-A3B) Continual pre-training of Qwen3 Swallow 30B-A3B (RL) on a mixture emphasizing medical-domain text (biomedical literature, medical synthetic data, medical QA, clinical guideline-style text) Swallow Project Apache 2.07
gpt-oss-20b-Ja-Fin
(CPT, Thinking)
2026 Finance GPT-OSS (20b) Continual pre-training on a Japanese financial corpus built from Common Crawl and other public sources (domain classification and quality filtering applied) Nomura Research Institute Apache 2.0
nekomata-14b-pfn-qfin
(qfin, qfin-inst-merge)
2024 Finance Qwen (14b) Continual pre-training on a finance-specific corpus (~8.1M documents, ~370M tokens): central bank press conferences and policy-meeting summaries, financial-institution reports, glossaries and corporate information, and Wikipedia finance articles (qfin-inst-merge is merged with the instruct model) Preferred Networks Tongyi Qianwen LICENSE
Qwen3-14B-Ja-Fin
(CPT, Thinking)
2026 Finance Qwen3 (14b) Continual pre-training on a Japanese financial corpus built from Common Crawl and other public sources (domain classification and quality filtering applied) Nomura Research Institute Apache 2.0
Watashiha-Llama-2-13B-Ogiri-sft
(sft, sft-neuron)
2024 Oogiri Llama 2 (13b) Pre-training: Japanese C4, CC-100, OSCAR, Japanese/English Wikipedia, proprietary data (65B tokens)
SFT: Ogiri data
Watashiha Llama 2 Community License
Medical-Qwen3-Swallow-8B 2026 Medicine Qwen3 (8b) Continual pre-training of Qwen3 Swallow 8B (RL) on a mixture emphasizing medical-domain text (biomedical literature, medical synthetic data, medical QA, clinical guideline-style text) Swallow Project Apache 2.07
MedExamDoc-Llama-3.1-Swallow-8B-Instruct-v0.5 2025 Medicine Llama 3.1 (8b) Fine-tuned with QLoRA on JMedBench and KokushiMD-10 Ingenta Llama 3.1 Community License
Llama 3.1 Future Code Ja 8B 2025 Coding Llama 3.1 (8b) Continual pre-training on The Stack v2 (code) and a subset of LLM-jp Corpus v3 (Japanese) — 204.9B code and 85.7B natural-language tokens Future Corp. Llama 3.1 Community License
Karamaru
(Karamaru-v1)
2025 Edo-period Japanese Llama 3 (8b) Continual pre-training on ~25M characters of Edo-period text (~13M human-transcribed: Minna de Honkoku, National Institute of Japanese Literature; ~12M AI-transcribed via the RURI kuzushiji OCR model) Sakana AI Llama 3 Community License
JPharmatron
(7B-base, 7B)
2025 Pharmaceutical Qwen2.5 (7b) Continual pre-training on Japanese pharmaceutical text (2B tokens) and English biomedical text (8B tokens) EQUES Inc. CC BY-SA 4.0
ELYZA-japanese-CodeLlama-7b
(7b, 7b-instruct)
2023 Coding Code Llama
(7b)
Pre-training: Japanese text from OSCAR, Wikipedia, and proprietary crawl data (18B tokens)
instruct: ELYZA's proprietary post-training (undisclosed)
ELYZA Llama 2 Community License
AIBunCho/japanese-novel-gpt-j-6b 2023 Storytelling GPT-J (6b) Additional training on Japanese novel data (the base Japanese GPT-J was pre-trained on CC100 Japanese, Wikipedia, and web data) Individual (Hiroyuki Osone) CreativeML OpenRAIL-M License
NovelAI/genji-jp 2022 Storytelling GPT-J (6b) Fine-tuned on a proprietary Japanese storytelling dataset NovelAI

Models built off non-Japanese LLMs (post-training only, no continual pre-training or details unknown)

General purpose

Release Year Base Model Training Data Developer License / Terms of Use
Rakuten AI 3.0
(RakutenAI-3.0)
2026 DeepSeek-V3 (671b) 15 Undisclosed Rakuten Apache 2.0
Llama 3.1 Shisa V2 405B
(405b)
2025 Llama 3.1 (405b) High-quality Japanese datasets with SFT/DPO Shisa.AI Llama 3.1 Community License
AXCXEPT/EZO-Qwen2.5-72B-Instruct
AXCXEPT/EZO-AutoCoTRAG-Qwen2.5-72B-Instruct_q4
2024 Qwen2.5 (72b) Axcxept Qwen License
ao-Karasu
(72B)
2024 Qwen1.5 (72b) ultra-orca-boros-en-ja-v1, OASST1, ShareGPT, Japanese technical blogs, News stories, QA site answers, undisclosed dataset Lightblue Tongyi Qianwen LICENSE (?)12
Shisa V2.1 70B
(70b)
2025 Llama 3.3 (70b) Combined SFT/DPO/RL/Model merging Shisa.AI Llama 3.3 Community License
shisa-ai/shisa-v2-llama3.3-70b 2025 Llama 3.3 (70b) Shisa.AI Llama 3.3 Community License
AXCXEPT/Llama-3.1-70B-EZO-1.1-it 2024 Llama 3.1 (70b) Axcxept Llama 3.1 Community License
Llama 3 shisa-v1-llama3-70b
(70b)
2024 Llama 3 (70b) ultra-orca-boros-en-ja-v1 Shisa.AI Llama 3 Community License (?)12
AIgroup-CVM-utokyohospital/Llama-2-70b-chat-4bit-japanese 2023 Llama 2 (70b) University of Tokyo Hospital Department of Cardiovascular Medicine AI Group Llama 2 Community License
doshisha-mil/llama-2-70b-chat-4bit-japanese-v1 2023 Llama 2 (70b) Doshisha University Media Informatics Lab
cyberagent/DeepSeek-R1-Distill-Qwen-32B-Japanese 2025 DeepSeek-R1-Distill-Qwen (32b) CyberAgent MIT
Flux-Japanese-Qwen2.5-32B-Instruct-V1.0
(V1.0)
2025 Qwen2.5-32B-Instruct (32b) Precise-tuning: Pinpointing Japanese knowledge, reasoning, and language circuits to apply adjustments to only 5% of parameters. Created 3 specialized models, then integrated through pinpoint merging FLUX Apache 2.0
karakuri-ai/karakuri-lm-32b-thinking-2501-exp 2025 QwQ (32b) KARAKURI Apache 2.0
shisa-ai/shisa-v2-qwen2.5-32b 2025 Qwen2.5 (32b) Shisa.AI Apache 2.0
AXCXEPT/EZO-Qwen2.5-32B-Instruct
AXCXEPT/EZO-AutoCoTRAG-Qwen2.5-32B-Instruct
2024 Qwen2.5 (32b) Axcxept Apache 2.0
cyberagent/DeepSeek-R1-Distill-Qwen-14B-Japanese 2025 DeepSeek-R1-Distill-Qwen (14b) CyberAgent MIT
Shisa V2.1 14B
(14b)
2025 Phi-4 (14b) Combined SFT/DPO/RL/Model merging Shisa.AI MIT
shisa-ai/shisa-v2-unphi4-14b 2025 Phi-4 (14b) Shisa.AI MIT
EZO-Phi-4
(phi-4-open-R1-Distill-EZOv1, phi-4-deepseek-R1K-RL-EZO)
2025 Phi-4 (14b) Axcxept MIT
Qarasu
(14B-chat-plus-unleashed)
2024 Qwen (14b) ultra-orca-boros-en-ja-v1, OASST1, ShareGPT, undisclosed dataset Lightblue Tongyi Qianwen LICENSE (?)12
Sparticle/llama-2-13b-chat-japanese-lora 2023 Llama 2 (13b) Sparticle
izumi-lab/llama-13b-japanese-lora-v0-1ep 2023 Llama (13b) University of Tokyo Izumi Lab
shisa-ai/shisa-v2-mistral-nemo-12b 2025 Mistral NeMo (12b) Shisa.AI Apache 2.0
AXCXEPT/EZO-Common-9B-gemma-2-it 2024 Gemma 2 (9b) Axcxept Gemma Terms of Use
AXCXEPT/EZO-Humanities-9B-gemma-2-it 2024 Gemma 2 (9b) Axcxept Gemma Terms of Use
CAT-Thinking
(8B)
2026 Qwen3 (8b) SFT (math/coding data generated by gpt-oss-120b and translated to Japanese via CAT-Translate-7b) / GRPO CyberAgent Apache 2.0
Shisa V2.1 8B
(8b)
2025 Qwen3 (8b) Combined SFT/DPO/RL/Model merging Shisa.AI Apache 2.0
AXCXEPT/Qwen3-EZO-8B-beta 2025 Qwen3 (8b) High-performance reasoning with Deep-Think technique Axcxept Apache 2.0
shisa-ai/shisa-v2-llama3.1-8b 2025 Llama 3.1 (8b) Shisa.AI Llama 3.1 Community License
AXCXEPT/Llama-3.1-8B-EZO-1.1-it 2024 Llama 3.1 (8b) Axcxept Llama 3.1 Community License
Llama 3 Suzume 8B
(8B-japanese, 8B-japanese-gguf)
2024 Llama 3 (8b) megagonlabs/instruction_ja, ShareGPT, undisclosed dataset Lightblue Llama 3 Community License (?)12
Llama 3 shisa-v1-llama3-8b
(8b)
2024 Llama 3 (8b) ultra-orca-boros-en-ja-v1 Shisa.AI Llama 3 Community License (?)12
AXCXEPT/Llama-3-EZO-8b-Common-it 2024 Llama 3 (8b) Axcxept Llama 3 Community License
lightblue/DeepSeek-R1-Distill-Qwen-7B-Japanese 2025 DeepSeek-R1-Distill-Qwen (7b) Lightblue Apache 2.0
ABEJA-Qwen2.5-7b-Japanese-v0.1
(v0.1)
2025 Qwen 2.5 (7b) ABEJA Apache 2.0
shisa-ai/shisa-v2-qwen2.5-7b 2025 Qwen 2.5 (7b) Shisa.AI Apache 2.0
Karasu DPO
(7B)
2025 Qwen 2.5 (7b) Lightblue Apache 2.0
ganchengguang/Yoko-7B-Japanese-v1 2023 Llama 2 (7b) Yokohama National University Mori Lab
Sparticle/llama-2-7b-chat-japanese-lora 2023 Llama 2 (7b) Sparticle
izumi-lab/llama-7b-japanese-lora-v0-5ep 2023 Llama (7b) University of Tokyo Izumi Lab
lightblue/jod 2023 Mistral-7B-SlimOrca (7b) Lightblue Apache 2.0
NTQAI/chatntq-7b-jpntuned 2023 RWKV-4 World (7b) NTQ Solution
Qwen3.5-FT-Japanese-CoT-4B 2026 Qwen3.5 (4b) Undisclosed Individual (Aname-Tommy) MIT
Borea
(Jp, Common, Coding)
2024 Phi-3.5 (3.8b) Axcxept MIT
Shisa V2.1 3B
(3b)
2025 Llama 3.2 (3b) Combined SFT/DPO/RL/Model merging Shisa.AI Llama 3.2 Community License
AXCXEPT/EZO-Llama-3.2-3B-Instruct-dpoE 2024 Llama 3.2 (3b) Axcxept Llama 3.2 Community License
Gemma-2-JPN
(2b-jpn-it)
2024 Gemma 2 (2b) Google Gemma Terms of Use
AXCXEPT/EZO-gemma-2-2b-jpn-it 2024 Gemma 2 (2b) Axcxept Gemma Terms of Use
AXCXEPT/EZO-Common-T2-2B-gemma-2-it 2024 Gemma 2 (2b) Axcxept Gemma Terms of Use
Shisa V2.1 1.2B
(1.2b)
2025 LFM2 (1.2b) Combined SFT/DPO/RL/Model merging Shisa.AI LFM Open License v1.0
LFM2.5-1.2B-JP
(1.2B-JP, 1.2B-JP-202606)
2026 LFM2.5 (1.2b) Undisclosed Liquid AI LFM Open License v1.0
Qwen3.5-FT-Japanese-CoT-0.8B 2026 Qwen3.5 (0.8b) Undisclosed Individual (Aname-Tommy) MIT

Domain specific

Release Year Domain Base Model Developer License
JMedLoRA
(llama2-jmedlora-6.89ep)
2023 Medicine Llama 2 (70b) University of Tokyo Hospital Department of Cardiovascular Medicine AI Group CC BY-NC 4.0
pfnet/Qwen3-1.7B-pfn-qfin 2025 Finance Qwen3 (1.72b) Preferred Networks PLaMo Community License
pfnet/Qwen2.5-1.5B-pfn-qfin 2025 Finance Qwen2.5 (1.54b) Preferred Networks PLaMo Community License

Merged models

Release Year Original Models (Japanese LLMs in bold) Developer License
EQUES/MedLLama3-JP-v2 2024 Llama 3 Swallow 8B (Instruct), OpenBioLLM-8B, MMed-Llama 3 8B, Llama 3 ELYZA JP 8B EQUES Llama 3 Community License
EvoLLM-JP-A
(v1-7B)
2024 Shisa Gamma 7B (v1), Arithmo2 Mistral 7B, Abel 7B 002 Sakana AI Apache 2.0
EvoLLM-JP
(v1-7B, v1-10B)
2024 Shisa Gamma 7B (v1), WizardMath-7B-V1.1, Abel 7B 002 Sakana AI MICROSOFT RESEARCH LICENSE
EQUES/TinyQwens-Merge-1.5B 2025 SakanaAI/TinySwallow-1.5B-Instruct, EQUES/TinySwallow-Stratos-1.5B, deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B, Qwen/Qwen2.5-1.5B-Instruct EQUES Inc. Apache 2.0

API-based models

Release Year Max Context Length Developer Platform
PLaMo API 2024 32,768 Preferred Networks self-owned
AI Novelist 2021 2,400 ~ 8,192 Bit192 self-owned
LHTM-OPT 2024 alt Inc. AWS Marketplace (SageMaker)
Syn
(Syn, Syn Pro)
2025 32,768 KARAKURI, Upstage AWS Marketplace (SageMaker)
tsuzumi
(tsuzumi-7b)
2024 NTT Microsoft Foundry

Encoder models

General purpose

Architecture Max Input Length Training Data Developer License HuggingFace? 16
ModernBERT-Ja ModernBERT 8,192 Japanese and English corpora SB Intuitions MIT ◯ (30m, 70m, 130m, 310m)
llm-jp-modernbert ModernBERT 8,192 Japanese subset of llm-jp-corpus-v4 (0.69T tokens) Research and Development Center for Large Language Models Apache 2.0
KyotoUniBERT BERT (base, large) 512 Japanese Wikipedia (18M articles) Kyoto University Language Media Processing Lab Apache 2.0
TohokuUniversityBERT BERT (base, large) 512 base (v1):
Japanese Wikipedia (17M articles / 2.6GB)
base (v2) & large:
Japanese Wikipedia 4.0GB
base (v3) & large (v2):
Japanese Wikipedia (4.9GB), Japanese CC‑100 (74.3GB)
Tohoku University NLP Group base (v1, v2) & large: CC BY‑SA 3.0
base (v3) & large (v2): Apache 2.0

(base (v1), base (v1, char-level), base (v2), base (v2, char-level), large, large (char-level), base (v3), base (v3, char-level), large (v2), large (v2, char-level))
TohokuNLP BERT-alpha 500M Llama-based encoder17 4,096
or
8,192
Japanese subset of llm-jp-corpus-v3 Tohoku University NLP Group Apache 2.0 ◯ (sq4096-alpha, sq8192-alpha)
ByBERT-JP Llama-based encoder17 100m, 200m, 400m: 3,072
v2-100m: 4,096
Japanese subset of llm-jp-corpus-v3
100m: 623B tokens
200m: 637B tokens
400m: 1.23T tokens
v2-100m: 2.76T tokens
Tohoku University NLP Group Apache 2.0 ◯ (100m, 200m, 400m, v2-100m)
NICT BERT BERT (base) 512 Japanese Wikipedia NICT CC BY 4.0
Laboro BERT BERT (base, large) 512 Japanese Web Corpus
(News and blogs, etc) (12GB)
Laboro.AI CC BY‑NC 4.0
colorfulscoop BERT BERT (base) 512 Japanese Wikipedia Colorful Scoop CC BY‑SA 3.0
UniversityOfTokyoBERT BERT (small) 512 Japanese Wikipedia (2.9GB) University of Tokyo Izumi Lab CC BY‑SA 4.0
chiTra (Sudachi Transformers) BERT (base) 512 NINJAL Web Japanese Corpus (148GB) NINJAL, WAP Tokushima Laboratory of AI and NLP Apache 2.0
ACCMS BERT BERT (base) 512 Japanese Wikipedia (3.3GB) Kyoto University ACCMS CC BY‑SA 4.0
HitachiBERT BERT (base) 512 Japanese Wikipedia, Japanese CC‑100 Hitachi CC BY‑NC‑SA 4.0 18
RetrievaBERT BERT 19 2,048 Japanese CommonCrawl, RefinedWeb, Chinese Wikipedia, Korean Wikipedia, The Stack Retrieva Apache 2.0
Bandai Namco DistilBERT DistilBERT 512 (Distillation of TohokuUniversityBERT(base)) Bandai Namco Research MIT
Laboro DistilBERT DistilBERT 512 (Distillation of Laboro BERT(base)) Laboro.AI CC BY‑NC 4.0
LINE DistilBERT DistilBERT 512 (Distillation of LINE internal BERT model) LINE Apache 2.0
rinna RoBERTa RoBERTa (base) 512 Japanese Wikipedia, Japanese CC‑100 rinna MIT
WasedaRoBERTa RoBERTa (base, large) 512 Japanese Wikipedia, Japanese CC‑100 Waseda Kawahara Lab CC BY‑SA 4.0
(base, large, large (seq512))20
InformatixRoBERTa RoBERTa (base) 512 Japanese Wikipedia, Web Articles
(25GB)
Informatix Apache 2.0
KyotoUniversityRoBERTa RoBERTa (base, large) 512 Japanese Wikipedia, Japanese CC‑100 Kyoto University Language Media Processing Lab CC BY‑SA 4.0
(base (char-level), large (char-level))
YokohamaNationalRoBERTa RoBERTa (base) 512 Japanese Wikipedia (3.45GB) Yokohama National University Mori Lab Apache 2.0
Megagon Labs RoBERTa RoBERTa (base)21 1,282 Japanese mC4 (200M sentences) Megagon Labs
(Recruit Co.,Ltd.)
MIT
ACCMS RoBERTa RoBERTa (base) 512 Japanese Wikipedia (3.3GB) + Japanese CC‑100 (70GB) Kyoto University ACCMS CC BY‑SA 4.0
CinnamonELECTRA ELECTRA (small) 512 Japanese Wikipedia Cinnamon Apache 2.0
Megagon Labs ELECTRA ELECTRA (base) 512 Japanese mC4 (200M sentences) Megagon Labs
(Recruit Co.,Ltd.)
MIT
UniversityOfTokyoELECTRA ELECTRA (small, base) 512 Japanese Wikipedia (2.9GB) University of Tokyo Izumi Lab CC BY‑SA 4.0
(small, base)
JapaneseRoFormer RoFormer (base) 512 Japanese Wikipedia (3.45GB) Yokohama National University Mori Lab Apache 2.0
JapaneseLUKE LUKE (base, large) 512 Japanese Wikipedia Studio Ousia Apache 2.0
(base, large)
KyotoUniversityDeBERTaV2 DeBERTaV2 (tiny, base, large) 512 Japanese Wikipedia, Japanese CC‑100, Japanese OSCAR
(171GB)
Kyoto University Language Media Processing Lab CC BY‑SA 4.0
(tiny, tiny (char-level), base, large)
KyotoUniversityDeBERTaV3 DeBERTaV3 (base) 512 llm-jp-corpus Kyoto University Language Media Processing Lab Apache 2.0
UniversityOfTokyoDeBERTaV2 DeBERTaV2 (small, base) 512 Japanese Wikipedia, Japanese Wikinews, Japanese CC-100, Japanese mC4, Japanese OSCAR University of Tokyo Izumi Lab CC BY-SA 4.0 ◯ (small, base)
GLOBIS DeBERTaV3 DeBERTaV3 (xsmall, base, large) 512 Wikipedia, WikiBooks, Aozora Bunko, Japanese CC-100, Japanese mC4, Japanese OSCAR GLOBIS CC BY-SA 4.0 ◯ (xsmall, base, large)
JapaneseBigBird BigBird (base) 4,096 Japanese Wikipedia, Japanese CC‑100, Japanese OSCAR Waseda Kawahara Lab CC BY‑SA 4.0
JapaneseLayoutLM LayoutLM (base) 512 Pre-trained on Japanese Wikipedia, initialized with TohokuUniversityBERT The Japan Research Institute, Limited CC BY-SA 3.0
layoutlmv3-japanese-preview LayoutLMv3 (base) 512 Initialized with WasedaRoBERTa (base), then pre-trained on ~20M Japanese web pages from NDL WARP and document images (PubLayNet, DocLayNet) Research and Development Center for Large Language Models Apache 2.0

Domain Specific

Domain Architecture Training Data Developer License HuggingFace?
JapaneseBlogELECTRA Colloquial language ELECTRA (small) Japanese Blog Corpus (354M sentences) Kitami Institute of Technology Masui-Ptaszynski Lab CC BY‑SA 4.0
JapaneseSpokenLanguageBERT Spoken language BERT (base) Additional training for TohokuUniversityBERT using Corpus of Spontaneous Japanese (CSJ)
(In the DAPT model, the diet record is also used)
Retrieva Apache 2.0
AcademicRoBERTa Science RoBERTa (base) CiNii Japanese Papers (6.3M sentences) Ehime University AI Lab Apache 2.0
local-politics-BERT Politics BERT (base) Wikipedia, Minutes of the National Diet, Minutes of the Local Assembly Japanese Local Assembly Minutes Corpus Project CC BY-SA 4.0 ◯ (SC-min, SC-minwiki, SC-2M-wiki, SC-2M-min, SC-2M-minwiki, FP-min, FP-minwiki) 22
UBKE-LUKE Economics LUKE (base) Japanese Wikipedia, Securities Reports, Economic News Articles Uzabase CC BY-NC
JapaneseFinancialBERT Finance BERT (small, base)23 Japanese Wikipedia, Japanese Financial Corpus (27M sentences/5.2GB) University of Tokyo Izumi Lab CC BY‑SA 4.0
(small, base)
JapaneseFinancialELECTRA Finance ELECTRA (small) Japanese Wikipedia (20M sentences/2.9GB), Japanese Financial Corpus (27M sentences/5.2GB) University of Tokyo Izumi Lab CC BY‑SA 4.0
JapaneseNewsBERT Business BERT (base) Japanese Business Articles (3M articles) Stockmark CC BY 4.0
JapaneseNewsXLNet Business XLNet (base) Japanese Business Articles (3M articles) Stockmark
※ Unofficial release
JapaneseNewsALBERT Business ALBERT (base) Japanese Business Articles (3M articles) Stockmark
MinpakuBERT Cultural Heritage BERT (base) Additional training with National Museum of Ethnology's cultural heritage data on top of Tohoku University BERT University of Hyogo Ohshima Lab MIT ◯ (minpaku-v1, minpaku-v3, minpaku-v3-no-additional-token)
JPharmaBERT Pharmacy BERT (base, large) Japanese Pharmaceutical Documents (2B tokens)
+ PubMed English Abstracts (8B tokens)
+ Multilingual Pharmaceutical Data (1.2B tokens)
EQUES Unknown ◯ (base, large)
medBERTjp Medicine BERT (base) Japanese Wikipedia, Japanese Medical Corpus ("今日の診療プレミアム/Today's Care Premium" Web Version) Osaka University Hospital
Medical Informatics Lab
CC BY‑NC‑SA 4.0
JMedRoBERTa Medicine RoBERTa (base) Japanese Medical Papers (11M sentences/1.8GB) NII Aizawa Lab CC BY‑NC‑SA 4.0
(ManbyoWordPiece, SentencePiece)24

Sentence and Document Embeddings 25

Bi-Encoders

Single-representation bi-encoders

Max Context Length Developer License
Ruri-v3
(v3-30m, v3-70m, v3-130m, v3-310m)
8,192 Nagoya University Sasano Group Apache 2.0
PLaMo-Embedding-1B
(1b)
4,096 Preferred Networks Apache 2.0
Sarashina-Embedding-v2
(v2-1b)
8,192 SB Intuitions Sarashina Model NonCommercial License
sbintuitions/sarashina-embedding-v1-1b 8,192 SB Intuitions Sarashina Model NonCommercial License
AMBER
(base, large)
512 Retrieva Apache 2.0
RoSEtta
(base-ja)
1,024 PKSHA Technology Apache 2.0
GLuCoSE v2
(base-ja-v2)
512 PKSHA Technology Apache 2.0
Ruri
(small, base, large, small-v2, base-v2, large-v2)
512 Nagoya University Sasano Group Apache 2.0
Japanese SimCSE
(unsup-simcse-ja-base, unsup-simcse-ja-large, sup-simcse-ja-base, sup-simcse-ja-large)
512 Nagoya University Sasano Group CC BY-SA 4.0
GLuCoSE
(base-ja)
512 PKSHA Technology Apache 2.0
colorfulscoop/sbert-base-ja Colorful Scoop CC BY‑SA 4.0
MU-Kindai/SBERT-JSNLI-base
MU-Kindai/SBERT-JSNLI-large
Kindai University
MU-Kindai/Japanese-SimCSE-BERT-base-unsup
MU-Kindai/Japanese-SimCSE-BERT-large-unsup
MU-Kindai/Japanese-SimCSE-RoBERTa-base-unsup
MU-Kindai/Japanese-SimCSE-BERT-base-sup
MU-Kindai/Japanese-SimCSE-BERT-large-sup
Kindai University MIT
pkshatech/simcse-ja-bert-base-clcmlp PKSHA Technology CC BY‑SA 4.0
MU-Kindai/Japanese-MixCSE-BERT-base
MU-Kindai/Japanese-MixCSE-BERT-large
Kindai University MIT
MU-Kindai/Japanese-DiffCSE-BERT-base Kindai University MIT
bclavie/fio-base-japanese-v0.1 Individual (Benjamin Clavié)
cl-nagoya/shioriha-large-pt Nagoya University Sasano Group

Multi-representation bi-encoders

Developer License
JaColBERTv2.5
(JaColBERTv2.4, JaColBERTv2.5)
Answer.AI MIT
JaColBERTv2
(JaColBERTv2)
Individual (Benjamin Clavié) MIT
JaColBERT
(JaColBERT)
Individual (Benjamin Clavié) MIT

Cross-Encoders

Developer License
Ruri-v3 Reranker
(310m)
Nagoya University Sasano Group Apache 2.0
Ruri-Reranker
(stage1-small, stage1-base, stage1-large, small, base, large)
Nagoya University Sasano Group Apache 2.0
hotchpotch/japanese-reranker-cross-encoder-xsmall-v1
hotchpotch/japanese-reranker-cross-encoder-small-v1
hotchpotch/japanese-reranker-cross-encoder-base-v1
hotchpotch/japanese-reranker-cross-encoder-large-v1
hotchpotch/japanese-bge-reranker-v2-m3-v1
Individual (Yuichi Tateno) MIT

Vision-Language Models

Text+Image to Text

Models built from scratch

*Includes VLMs newly built by combining an LLM with a vision encoder (the base LLM can be either domestic or non-Japanese).

General purpose
Release Year Architecture Training Data Developer License / Terms of Use
Stockmark-2-VL-100B-beta
(100B-beta)
2025 LLaVA-OneVision 3-stage training: alignment pre-training, caption expansion, instruction and reasoning fine-tuning
Synthetic data: Generated from Qwen2.5-VL-72B
Stockmark Qwen License
Llama-3.1-70B-Instruct-multimodal-JP-Graph
(v0.1)
2025 LLaVA (Llama-3.1-Swallow-70B-Instruct-v0.3 + Qwen2-VL-7B-Instruct) Over 6 million synthetic visual data specialized for chart and graph understanding (text, pie charts, bar charts, flowcharts, etc.), real data (with FastLabel collaboration) Ricoh Llama 3.1 Community License & Gemma Terms of Use & Qwen License & MIT & Apache 2.0
KARAKURI VL
(32b-instruct-2507, 32b-thinking-2507-exp)
2025 Vision-Language (based on Qwen2.5-VL-32B) Custom dataset specialized for Japanese computer use: Japanese computer operation records, Japanese document image QA, visual information interpretation, Japanese OCR, flowchart comprehension
3-stage training: Supervised Fine-Tuning (SFT) + model merging + reinforcement learning
*thinking model shows reasoning process explicitly using Chain of Thought (CoT) approach
KARAKURI Apache 2.0
Heron-NVILA
(1B, 2B, 15B, 33B)
2025 NVILA 3-stage training: Alignment (558k Japanese image-text pairs + 595k LLaVA-Pretrain), Pre-training (MOMIJI 13M, Japanese image-text pairs 6M, Japanese interleaved data 2M, coyo-700m 6M, mmc4-core 4M, Wikipedia-ja, LLaVA-Pretrain-JA, STAIR captions), Supervised fine-tuning (LLaVA-instruct-v1.5-en, LLaVA-instruct-ja, Japanese photos conversation, JA-VG-VQA conversation, SynthDog-ja, AI2D, SynthDog-en, Sherlock) Turing Apache 2.0 & OpenAI Terms of Use
NABLA-VL
(15B)
2025 microsoft/phi-4 + HuggingFaceM4/siglip-so400m-14-980-flash-attn2-navit Supports single image, multiple images, and video inputs. Training details undisclosed NABLAS Apache 2.0
Sarashina2-Vision
(8b, 14b)
2025 Sarashina2 + Qwen2-VL + 2-layer MLP 3-stage training: Projector warmup (LLaVA-Pretrain 78M English tokens), Vision encoder pre-training (CC3M, CC12M, llm-jp-japanese-image-text-pairs, internal OCR dataset, internal chart caption synthetic dataset 3.8B Japanese + 7.7B English tokens), Visual instruction tuning (Japanese Visual Genome VQA, OCR-VQA, TextVQA, PlotQA, CLEVR translated, DOCCI translated, internal datasets 2.5B Japanese + 1.0B English tokens) SB Intuitions MIT
Asagi
(2B, 4B, 8B, 14B)
2025 LLaVA Newly crawled Japanese website images, existing Japanese datasets, Japanese translations of English datasets ~20M samples (data synthesis using English VLM Phi-3.5-vision-instruct and Japanese LLM CALM3-22B-Chat) University of Tokyo Machine Intelligence Lab. Apache 2.0
llava-calm2-siglip
(llava-calm2-siglip)
2024 LLaVA coversational data generated from MS-COCO and VisualGenome CyberAgent Apache 2.0
LLM-jp-3 VILA 14B
(14b)
2024 LLaVA Japanese image text pairs, LLaVA-Pretrain, Japanese interleaved data, coyo (subset), mmc4-core (subset), llava-instruct-ja, japanese-photos-conv, ja-vg-vqa, synthdog-ja, LLaVA-1.5 instruction data (subset) Research and Development Center for Large Language Models Apache 2.0 & OpenAI Terms of Use
LLM-jp-4-VL 9B beta
(9b-beta)
2026 LLaVA (llm-jp-4-8b-instruct + siglip2-so400m-patch16-512 + 2-layer MLP) Jagle and others, ~33.4M samples / ~180B tokens (Japanese + English) Research and Development Center for Large Language Models Apache 2.0
PLaMo 2.1 8B VL
(8b-vl)
2026 LLaVA (PLaMo 2.1-8B + siglip2-so400m-patch14-384 + MLP) 2-stage training: Stage 1.0: Image adapter training (web-scale image-caption data with Japanese:English = 75:25), Stage 1.5: Instruction Tuning and Japanese visual adaptation via LoRA (VQA, Visual Grounding, object detection, anomaly detection, factory work understanding) Preferred Networks PLaMo community license
Heron
(blip-ja-stablelm-base-7b-v0, blip-ja-stablelm-base-7b-v1, blip-ja-stablelm-base-7b-v1-llava-620k, git-ja-stablelm-base-7b-v0, git-ELYZA-fast-7b-v0, git-ja-stablelm-base-7b-v1)
2023 BLIP-2 / GIT v1: LLaVA-Instruct-150K-JA or LLaVA-Instruct-620K-JA
v0: LLaVA-Instruct-150K-JA, Japanese STAIR Captions, Japanese Visual Genome VQA dataset
Turing CC BY-NC 4.0
Japanese Stable VLM
(japanese-stable-vlm)
2023 LLaVA Japanese CC12M, STAIR Captions, Japanese Visual Genome VQA dataset Stability AI STABILITY AI JAPANESE STABLE VLM COMMUNITY LICENSE
Japanese InstructBLIP Alpha
(japanese-instructblip-alpha)
2023 InstructBLIP Japanese CC12M, STAIR Captions, Japanese Visual Genome VQA dataset Stability AI JAPANESE STABLELM RESEARCH LICENSE
rinna MiniGPT-4
(bilingual-gpt-neox-4b-minigpt4)
2023 MiniGPT-4 CC12M, COCO 2014, Visual Genome, STAIR Captions, Japanese Visual Genome VQA dataset rinna MIT
Sarashina2.2-Vision-3B
(3.8b)
2025 Sarashina2.2-3B-Instruct + SigLIP + 2-layer MLP 4-stage training + Post-training: Projector warmup (English image captions), Vision encoder pre-training (Japanese charts, OCR, captions), Full model pre-training (interleaved image-text data), Supervised fine-tuning
Post-training: Mixed Preference Optimization
(Total: 103B Japanese + 157.1B English tokens)
SB Intuitions MIT
Jagle-VL 2.2B
(2.2b-Jagle, 2.2b-FineVision, 2.2b-Jagle-FineVision)
2026 InternVL3 (Qwen3-1.7B + siglip2-so400m-patch16-512 + 2-layer MLP) Trained on Jagle (large-scale Japanese multimodal post-training dataset, ~9.2M samples) and FineVision (model name suffixes -Jagle / -FineVision / -Jagle-FineVision indicate which combination is used) Research and Development Center for Large Language Models Apache 2.0
PLaMo 2.1 2B VL
(2b-vl)
2026 LLaVA (PLaMo 2.1-2B + siglip2-so400m-patch14-384 + MLP) 2-stage training: Stage 1.0: Image adapter training (web-scale image-caption data with Japanese:English = 75:25), Stage 1.5: Instruction Tuning and Japanese visual adaptation via LoRA (VQA, Visual Grounding, object detection, anomaly detection, factory work understanding) Preferred Networks PLaMo community license
Domain Specific
Release Year Architecture Domain Developer License
Med-Asagi
(14b-reasoning_beta)
2026 LLaVA Medicine University of Tokyo Machine Intelligence Lab. CC BY-SA 4.0
watashiha/Watashiha-Llama-2-13B-Ogiri-sft-vlm 2024 LLaVA Oogiri Watashiha Llama 2 Community License

Models built off non-Japanese VLMs

Release Year Base Model Training Data Developer License
Stockmark-DocReasoner-Qwen2.5-VL-32B
(32B)
2026 Qwen2.5-VL-32B-Instruct 2-stage curriculum learning (SFT): Stage 1: ~1.1M samples (incl. 600K synthetic) for basic document structure understanding, Stage 2: ~1.4M samples (incl. 500K synthetic) for multi-step reasoning and Chain-of-Thought
Persona-based synthetic data (leveraging Nemotron-Personas-Japan), rendering with matplotlib/HTML/Plotly/LaTeX/mermaid, quality verification via VLM-as-a-judge
Stockmark Apache 2.0
AXCXEPT/EZO-InternVL2-26B 2024 InternVL2 -  Axcxept MIT
KARAKURI VL 2
(8b-thinking-2603)
2026 Qwen3-VL-8B-Thinking Undisclosed KARAKURI Apache 2.0
Qwen-3-VL-Ricoh-8B-20260227
(8B-20260227)
2026 Qwen3-VL-8B-Thinking Reasoning process via Reinforcement Learning (RL) Ricoh Apache 2.0

Merged models

Release Year Original Models (Japanese LLMs in bold) Developer License
Llama-3-EvoVLM-JP-v2
(v2)
2024 Mantis-8B-SigLIP-Llama-3, Llama-3-ELYZA-JP-8B, Bunny-v1.1-Llama-3-8B-V Sakana AI Llama 3 Community License
AXCXEPT/Llama-3-EZO-VLM-1 2024 - (trained from Llama-3-EvoVLM-JP-v2) Axcxept Llama 3 Community License
EvoVLM-JP
(v1-7B)
2024 Shisa Gamma 7B (v1), LLaVA-1.6-Mistral-7B Sakana AI Apache 2.0

Text to Image

General Purpose

Release Year Architecture Training Data Developer License
CommonArt β
(commonart-beta)
2024 PixArt-Σ CommonCatalog-cc-by, Megalith-10M, Smithonian Open Access, ArtBench (CC-0 only) AI Picasso Apache 2.0
EvoSDXL-JP
(v1)
2024 Stable Diffusion - (merged from several diffusion models, including Japanese Stable Diffusion XL) Sakana AI Apache 2.026
Japanese Stable Diffusion XL
(japanese-stable-diffusion-xl)
2023 Stable Diffusion undisclosed Stability AI STABILITY AI JAPANESE STABLE DIFFUSION XL COMMUNITY LICENSE
TohokuUniversity Stable Diffusion
(base, refiner)
2023 Stable Diffusion WMT2023 Shared Task English-Japanese parallel corpus, about 13 million captions from laion2B-multi Tohoku University NLP Group CreativeML OpenRAIL-M License
rinna Stable Diffusion
(japanese-stable-diffusion)
2022 Stable Diffusion LAION-5B Japanese Subset (100M images) rinna CreativeML OpenRAIL-M License

Domain Specific

Release Year Architecture Domain Developer License
Evo-Nishikie
(v1)
2024 Stable Diffusion (ControlNet) Ukiyo-e Sakana AI Apache 2.026
Evo-Ukiyoe
(v1)
2024 Stable Diffusion Ukiyo-e Sakana AI Apache 2.026

Text to Video

Release Year Architecture Training Data Developer License
AIdeaLab VideoJP
(AIdeaLab-VideoJP)
2025 CogVideoX Pixabay, FineVideo AIdeaLab Apache 2.0

Others

Release Year Architecture Training Data Developer License
llm-jp-clip
(llm-jp-clip-vit-base-patch16, llm-jp-clip-vit-large-patch14)
2024 CLIP Translation of about 1.5 billion captions from the English subset of ReLAION-5B Research and Development Center for Large Language Models Apache 2.0
LY CLIP
(clip-japanese-base, v2)
2025 CLIP CommonCrawl, CC12M, YFCC100M
(v2: ~2B image-text pairs from Common Crawl + knowledge distillation)
LY Corp. Apache 2.0
Recruit CLIP
(japanese-clip-vit-b-32-roberta-base)
2023 CLIP about 120 million captions from laion2B-multi Recruit Co.,Ltd. CC BY-4.0
Japanese Stable CLIP
(japanese-stable-clip-vit-l-16)
2023 SigLIP CC12M translated to Japanese, STAIR Captions Stability AI STABILITY AI JAPANESE STABLE CLIP COMMUNITY LICENSE
rinna CLIP
(japanese-clip-vit-b-16)
2022 CLIP CC12M translated to Japanese rinna Apache 2.0
rinna CLOOB
(japanese-cloob-vit-b-16)
2022 CLOOB CC12M translated to Japanese rinna Apache 2.0
HAKUHODO Technologies CLIP
(base, deeper, wider)
2024 CLIP about 120 million captions from laion2B-multi HAKUHODO Technologies CC BY-NC-SA 4.0

Speech-Language Models

Automatic Speech Recognition

Release Year Architecture Training Data Developer License
Nue ASR
(nue-asr)
2023 Nue ASR
(HuBERT + LLM)
ReazonSpeech rinna Apache 2.0
Kana-Whisper27
(kana-whisper)
2026 Whisper (large-v3-turbo) Corpus of Spontaneous Japanese (CSJ) SB Intuitions MIT
Kotoba-Whisper
(v1.0, v1.0-ggml, v1.0-faster, v1.1, bilingual-v1.0, bilingual-v1.0-ggml, bilingual-v1.0-faster, v2.0, v2.0-ggml, v2.0-faster, v2.1, v2.2)
2024 Distil-Whisper ReazonSpeech Kotoba Technologies Apache 2.0
ReazonSpeech
(espnet-v1, espnet-next, espnet-v2, nemo-v2)
2024 ESPnet (Conformer-Transducer) / NeMo (FastConformer-RNNT) ReazonSpeech Reazon Holdings Apache 2.0
Reazon HuBERT ASR
(rs35kh, rs35kh-bpe)
2025 HuBERT ReazonSpeech v2.0 Reazon Holdings Apache 2.0
Reazon Zipformer ASR
(rs35kh, rs35kh-bpe)
2025 Zipformer ReazonSpeech v2.0 Reazon Holdings Apache 2.0
Reazon wav2vec 2.0 ASR
(base-rs35kh, large-rs35kh)
2024 wav2vec 2.0 ReazonSpeech v2.0 Reazon Holdings Apache 2.0

Text-to-Speech (TTS)

Release Year Architecture Training Data Developer License
Sarashina2.2-TTS
(sarashina2.2-tts)
2026 TTS built on Sarashina2.2 (0.5B)
(CosyVoice + HiFT-GAN)
Legally obtained audio data (purchased sources, public speech archives, and data collected in compliance with applicable domestic laws) SB Intuitions Sarashina Model NonCommercial License
Kotoba-Speech
(v0.1)
2024 Transformer undisclosed Kotoba Technologies Apache 2.0

Speech Foundation Models / Spoken Dialogue

Release Year Architecture Training Data Developer License
LLM-jp-Moshi-v1
(llm-jp-moshi-v1)
2026 Transformer-based text and speech foundation model (Moshi) J-CHAT (~69,000 hours), LLM-jp-Zoom1 (~1,000 hours) 大規模言語モデル研究開発センター Apache 2.0
J-Moshi
(j-moshi, j-moshi-ext)
2025 Transformer-based text and speech foundation model (Moshi) Speech dialogue corpus (J-CHAT, Japanese Callhome, CSJ, travel agency dialogue corpus, proprietary chat dialogue corpus, proprietary consultation dialogue corpus), text dialogue corpus (Japanese PersonaChat, Japanese EmpatheticDialogues, Japanese daily dialogue corpus, RealPersonaChat) Nagoya University Higashinaka Lab CC BY-NC 4.0
LFM2.5-Audio-1.5B-JP
(1.5B-JP)
2026 LFM2.5-Audio
(LFM2 + FastConformer)
Undisclosed Liquid AI LFM Open License v1.0

Feature Extraction

Release Year Architecture Training Data Developer License
NEST-Ja
(0.1b, 0.6b)
2026 NEST (FastConformer) ReazonSpeech v2.0 SB Intuitions MIT
Kushinada
(base, large)
2025 HuBERT 60k hours of audio extracted from large-scale Japanese TV broadcast audio data Intelligent Media Processing Research Team, AIST Apache 2.0
Reazon HuBERT
(base-k2)
2025 HuBERT ReazonSpeech Reazon Holdings Apache 2.0
UniversityOfTokyoHuBERT
(base-jtube)
2024 HuBERT JTubeSpeech University of Tokyo
Saruwatari & Takamichi Lab
MIT
rinna HuBERT
(base, large)
2023 HuBERT ReazonSpeech rinna Apache 2.0
Izanami
(base, large)
2025 wav2vec 2.0 60k hours of audio extracted from large-scale Japanese TV broadcast audio data Intelligent Media Processing Research Team, AIST Apache 2.0
Reazon wav2vec 2.0
(base, large)
2024 wav2vec 2.0 ReazonSpeech Reazon Holdings Apache 2.0
rinna wav2vec 2.0
(base)
2024 wav2vec 2.0 ReazonSpeech rinna Apache 2.0
rinna data2vec Audio
(base)
2024 data2vec Audio ReazonSpeech rinna Apache 2.0
Reazon Zipformer
(base-k2)
2025 Zipformer ReazonSpeech Reazon Holdings Apache 2.0

Music-Language Models

Music-Text Conversion

Release Year Architecture Training Data Developer License
Japanese MULAN
(japanese-mulan-base)
2025 MULAN (AST + GLuCoSE) ~20k internal music-text pairs LY Corporation Apache 2.0

Evaluation Benchmarks for Japanese LLMs

Hybrid Benchmarks

Description Developer
Nejumi LLM Leaderboard4 Comprehensively evaluates the Japanese language capabilities of LLMs across application development (coding, function calling), reasoning abilities (mathematical, logical, abstract reasoning), specialized knowledge, and safety evaluation (instruction following, hallucination suppression). The introduction of high-difficulty benchmarks clarifies performance differences among top-tier models. For more details, see this article. Weights & Biases
Swallow LLM Leaderboard v2 Conducts a comprehensive evaluation of various LLMs based on three types of tasks: Japanese language understanding and generation tasks, Japanese multi-turn dialogue tasks, and English language understanding and generation tasks. v2 supports reasoning-focused models by adopting zero-shot inference and chain-of-thought prompting, evaluating on more challenging benchmarks (12 total tasks: 6 Japanese, 6 English). Also publishes swallow-evaluation, an evaluation script that integrates and improves existing LLM evaluation tools, plus the newly released swallow-evaluation-instruct for reasoning-type models. Swallow Project

Traditional Benchmarks based on Natural Language Understanding tasks

Description Developer
Open Japanese LLM Leaderboard Evaluates Japanese language models across 14 categories with 71+ tasks using llm-jp-eval. LLM-jp, Hugging Face
llm-jp-eval A tool that evaluates Japanese LLMs automatically across multiple datasets.
The complete list of supported datasets can be found here (which also includes tasks such as JNLI and JCommonsenseQA from JGLUE).
LLM-jp
JP Language Model Evaluation Harness A fork by Stability AI of EleutherAI/lm-evaluation-harness. It is a tool for automatically evaluating Japanese LLMs across multiple datasets.
The complete list of supported datasets can be found here (which also includes tasks such as JNLI and JCommonsenseQA from JGLUE).
Stability AI
JGLUE Japanese version of the GLUE benchmark suite, including the MARC-ja, JCoLA, JSTS, JNLI, JSQuAD, and JCommonsenseQA tasks. JCoLA is by the University of Tokyo's Oseki Lab. See here and here (ja only) for further details about each task. Waseda University Kawahara Lab and Yahoo
JMMLU A benchmark constructed as a Japanese version of the MMLU Benchmark, consisting of multiple-choice questions from a wide range of academic fields including natural sciences, humanities, and social sciences. In addition to translating the original MMLU, it features newly added problems based on the unique cultural background of Japan (Japan-specific problems). Waseda University Kawahara Lab

Benchmarks on open-ended generative tasks

Description Developer
llm-jp-judge An integrated evaluation tool for Japanese LLMs using the LLM-as-a-Judge approach. It evaluates across four categories: Japanese quality (rating accuracy, fluency, detail, and relevance on a 1-5 scale), Japanese safety, MT-Bench (English), and MT-Bench (Japanese). The tool separates generation and evaluation phases and supports multiple inference clients including vLLM, OpenAI API, Azure OpenAI, and AWS Bedrock. Details here. Research and Development Center for Large Language Models
Japanese MT-bench The Japanese version of MT-bench asks about multi-turn conversational ability. It includes 80 questions, 10 each, from 8 categories: Writing, Roleplay, Reasoning, Math, Coding, Extraction, STEM, Humanities. Some questions have been modified to fit with Japanese culture during the production of the Japanese version. It also includes a script that performs a 10-level absolute evaluation by GPT-4. Stability AI
ELYZA-tasks-100 Ranking based on model responses to 100 complex and diverse tasks, including tasks testing summarization, correction, abstraction, induction, and other skills. Uses humans to score the model responses and then ranks models based on their mean scores. ELYZA
Preferred Generation Benchmark
(pfgen-bench)
A benchmark to measure the Japanese language generation ability of LLMs based on 50 common sense questions unique to the Japanese context. It evaluates along three axes: Fluency, Truthfulness, and Helpfulness. The evaluation is conducted without using LLM-as-a-Judge by calculating n-gram or rule-based metrics. Preferred Elements (Preferred Networks)
Rakuda Benchmark Ranking based on model answers to 40 open-ended questions on Japanese geography, history, politics, and society. Uses GPT-4 to judge model outputs pairwise, and then ranks models by fitting a Maximum Likelihood Elo/Bradley-Terry model to GPT-4's preferences. YuzuAI
Japanese Vicuna QA Benchmark This is the Japanese version of vicuna-blog-eval, which is the predecessor of MT-Bench. It includes 80 questions on general knowledge, role-playing, common sense, Fermi estimation, counterfactual thinking, coding, mathematics, and writing. It also includes a script for automatic evaluation by GPT-4 (win-rate calculation). The leaderboard can be found here. Kyoto University Language Media Processing Lab
Tengu-Bench Includes 120 free-form questions from various categories. Categories of questions: table interpretation, logic puzzles, idea generation, function calling, long document summarization (over a thousand tokens), conversation summarization, long document closed QA (over a thousand tokens), honorifics, project creation, math, translation, extraction, ethical control, cost estimation, Japan, chit-chat, puns, formatting, construction, business, legal judgment, politics, hypothetical questions. Lightblue
Shaberi A framework that can collectively evaluate the Japanese MT-bench, Rakuda Benchmark, ELYZA-tasks-100, and Tengu-Bench. There is also a fork by Shisa.AI. Lightblue

Benchmarks for measuring performance in specific domains

Description Developer
Japanese Language Model Financial Evaluation Harness A benchmark for Japanese LLM in the financial sector. It includes tasks such as sentiment analysis in finance (chabsa), basic knowledge tasks in securities analysis (cma_basics), tasks related to audits in certified public accountant examinations (cpa_audit), multiple choice question tasks in financial planner exams (fp2), and mock exam tasks for securities salespeople exams (security_sales_1). For more details, please see here. Preferred Networks
pfmt-bench-fin-ja A benchmark for measuring the generation capabilities of Japanese LLMs in the financial domain. Preferred Networks
jfinqa A Japanese financial numerical reasoning QA benchmark. Contains 1,000 numerical reasoning questions extracted from securities reports of 68 companies. Evaluates financial reasoning capabilities including arithmetic operations, ratio calculations, and DuPont analysis. Available on PyPI and HuggingFace. Individual (ajtgjmdjp)
Stockmark Business Questions The collection includes 50 questions that probe knowledge on topics such as market trends, current affairs, social issues, and business trends. Stockmark
JMED-LLM A dataset for evaluating LLMs in the Japanese medical domain. It compiles previously developed Japanese medical language processing tasks for LLM benchmarking. NAIST Social Computing Lab.
JMedBench A benchmark for LLMs in the Japanese medical field. It includes 20 datasets in 5 types of tasks: multi-choice question-answering, machine translation, named entity recognition, document classification, and semantic textual similarity (some datasets are borrowed from JMMLU and JMED-LLM). A tool called med-eval is developed to facilitate evaluation on JMedBench. NII Aizawa Lab
Japanese Medical Language Model Evaluation Harness A benchmark for evaluating Japanese LLMs in the medical domain in both Japanese and English, executable by a single command. Individual (Issey Sukeda)
YakugakuQA A Japanese pharmaceutical domain evaluation dataset based on national pharmacist licensing exams. Tests factual pharmaceutical knowledge. EQUES Inc.
NayoseQA A Japanese pharmaceutical domain evaluation dataset for cross-lingual terminology normalization. Tests understanding of synonyms and technical terms. EQUES Inc.
SogoCheck A novel task designed to assess consistency reasoning between paired statements. A challenging reasoning task where even GPT-4o performs poorly. EQUES Inc.
MedRECT A benchmark for evaluating the ability to detect and correct medical errors in clinical records. It consists of three tasks: error detection, error sentence identification, and error correction. It includes a Japanese version (663 samples) and an English version (458 samples), with the Japanese version constructed based on the Japanese Medical Licensing Examination. Preferred Networks
karakuri-bench A dataset for measuring performance of Japanese LLMs in customer support. KARAKURI

Benchmarks for measuring factuality and safety

Description Developer
JTruthfulQA The Japanese version of the dataset for evaluating the factuality of LLMs TruthfulQA. It includes questions about superstitions and other beliefs held by some people that are not factual, as well as questions about Japan-specific knowledge, all collected from scratch. Waseda University Kawahara Lab
JCommonsenseMorality A dataset on Japanese commonsense morality. Sentences describing actions are labeled with binary values indicating whether they are morally wrong or acceptable. Hokkaido University Language Media Lab
JBBQ The Japanese version of the social bias QA dataset BBQ, developed through translation, revision, and addition of questions based on Japanese culture and customs. University of Tokyo Yanaka Lab

Benchmarks for measuring logical reasoning capabilities

Description Developer
JFLD (Japanese Formal Logic Deduction) A dataset for evaluating deductive reasoning capabilities of Japanese LLMs (the Japanese version of the FLD (Formal Logic Deduction) proposed by the same authors). It is characterized by being composed of counterfactual samples to evaluate apart from the knowledge the LLM possesses. Hitachi
JHumanEval A Japanese version of the HumanEval benchmark, which assesses the ability to generate Python code from English instructions. In creating the Japanese version, the text was first machine-translated and then manually corrected. Japan Women's University Kuramitsu Lab
JMultiPL-E A dataset for evaluating code generation capabilities across 17 programming languages (C++, C#, Go, Java, JavaScript, PHP, Ruby, Rust, Scala, Swift, TypeScript, etc.) based on OpenAI HumanEval. Measures multilingual code understanding and generation performance. Tohoku University Natural Language Processing Group

Benchmarks on instruction-following ability

Description Developer
LCTG Bench A benchmark for the controllability of Japanese LLMs. It evaluates whether LLMs can adhere to constraints in four aspects: output format, character count, keywords, and forbidden words. The quality of the generated text is also evaluated. CyberAgent
JFBench A benchmark for evaluating the instruction-following ability of Japanese LLMs. In addition to 6 constraint groups translated from IFBench, 10 new groups specific to Japanese were created (e.g., polite/plain style, mixed hiragana/katakana/kanji, numerical notation). It has 16 constraint groups and 174 constraint types, evaluating a total of 1,600 samples with constraint counts of 1/2/4/8. Preferred Networks

Benchmarks for embedding models

Description Developer
JMTEB A benchmark developed as the Japanese version of MTEB. It consists of tasks such as document clustering, text classification, sentence similarity, sentence pair labeling prediction, and text extraction (a reranking task was recently added). SB Intuitions
JQaRA A dataset for evaluating Japanese document extraction and reranking accuracy. Each of the 1,667 questions is assigned 100 candidate documents, of which at least one can answer the question. The questions are taken from JAQKET, and the candidate documents are sourced from Japanese Wikipedia. Individual (Yuichi Tateno)
JaCWIR A dataset created for evaluating document extraction and reranking in domains other than Wikipedia. Each of the 5,000 questions is assigned one Web page that serves as the source of the question and 99 unrelated Web pages. Individual (Yuichi Tateno)

Benchmarks for vision-language models

Description Developer
llm-jp-eval-mm A framework for evaluating Japanese VLMs. It supports major Japanese tasks including Japanese-Heron-Bench, JA-VLM-Bench-In-the-Wild, JA-Multi-Image-VQA, JDocQA, JMMMU, JIC-VQA, MECHA-ja, CC-OCR, and CVQA, and provides LLM-as-a-Judge automatic scoring, fast batch inference via vLLM, and a web dashboard for inspecting results. Research and Development Center for Large Language Models
simple-evals-mm A framework for evaluating Japanese VLMs on 20 tasks (10 Japanese tasks + 10 English tasks). Japanese tasks include JAMMEval, BusinessSlideVQA, JMMMU, and MECHA-ja, while English tasks include AI2D, ChartQA, MMMU, and others. Research and Development Center for Large Language Models
JAMMEval A collection of evaluation datasets for Japanese VLMs, built by refining seven existing Japanese multimodal benchmarks (CC-OCR-JA, CVQA-JA, Heron-Bench, JA-Multi-Image-VQA, JA-VLM-Bench, JDocQA, JGraphQA). Research and Development Center for Large Language Models
BusinessSlideVQA A question-answer dataset with 220 questions about complex Japanese business slide images, designed to evaluate document comprehension capabilities. Stockmark
JA-Business-Doc-RQ-Bench A benchmark for evaluating multi-step reasoning ability on Japanese business documents. Consists of 229 questions (Yes/No, Factoid, and Numerical) requiring 3-5 reasoning steps over 4 image types (Chart, Table, Diagram, and Document). All images are synthetic, with questions and answers manually written. Stockmark
JMMMU A benchmark constructed as the Japanese version of MMMU Benchmark. It consists of 720 translated MMMU problems and 600 new problems unique to Japanese culture. University of Tokyo Aizawa Lab
JDocQA A question-answer dataset based on Japanese documents (pamphlets, slides, reports, websites), consisting of a total of 11,600 questions. It includes various question formats, including unanswerable questions. NAIST Watanabe Lab
JGraphQA A benchmark for evaluating Japanese chart comprehension. It consists of 100 images of four chart types (pie charts, line charts, bar charts, and tables) collected from Japanese corporate IR materials, with 2 manually created and verified question-answer pairs per image, for a total of 200 questions. Ricoh
Heron VLM Leaderboard powered by Nejumi/WandB Summarizes the evaluation results of Japanese-Heron-Bench and LLaVA-Bench-In-the-Wild (Japanese). Turing, Weights & Biases
Japanese-Heron-Bench 21 images are assigned a total of 102 questions. It is characterized by image-question pairs that require knowledge related to Japan. Turing
JA-VLM-Bench-In-the-Wild A dataset independently prepared by Sakana AI to evaluate EvoVLM-JP-v1-7B. It consists of 50 questions assigned to 42 images. It is characterized by images and questions that require knowledge about Japan. Sakana AI
JA-Multi-Image-VQA A dataset for evaluating the question-answering ability in Japanese for multiple images. Sakana AI
LLaVA-Bench-In-the-Wild (Japanese) This is the Japanese version of LLaVA-Bench-In-the-Wild, translated using DeepL. It consists of 60 questions assigned to 24 images. Turing
LLaVA-Bench (COCO) Japanese This is the Japanese version, translated by DeepL, of the LLaVA-Bench (COCO) dataset used to evaluate LLaVA. It consists of 30 images, each with 3 types of questions assigned to them. Turing
Japanese Visual Genome VQA dataset A question-and-answer dataset annotated based on images from the Visual Genome dataset. A subset of this dataset, JA-VG-VQA-500, consisting of 500 questions, is sometimes used as a benchmark for evaluating VLMs. Yahoo
japanese-bizform-table-kie A benchmark for evaluating the accuracy of key information extraction from non-standard business forms. It consists of 50 form types with a total of 2,500 document images. AI inside

Benchmarks/Datasets for Speech-Language Models

Description Developer
VoiceBench-ja A benchmark that evaluates Japanese speech-language models (audio-input LLMs) by measuring the performance gap between audio input and text input. It consists of four subsets: Elyza (36 questions derived from ELYZA-tasks-100), Spoken-Elyza (34 questions refined for spoken dialogue), M-IFEval (172 instruction-following questions), and JamC-QA (1,452 multiple-choice questions on Japan-specific knowledge). The audio was synthesized with SB Intuitions' in-house TTS using speaker prompts from the JVS corpus (text is CC BY-SA 4.0; audio is not for commercial use or redistribution). SB Intuitions
Joyo Kanji Yomi Benchmark A benchmark for evaluating the kanji-level reading accuracy of Japanese TTS. It consists of 13,095 sentences covering all 2,136 Joyo kanji and their 4,378 readings, with each sentence written so that the context uniquely determines the target reading. Every sentence carries a full katakana reading annotation in which the substring corresponding to the target kanji is delimited, and all sentences were verified by 35 native Japanese speakers through a three-stage review. An evaluation toolkit that transcribes synthesized speech into katakana with Kana-Whisper and compares it against the reference reading is also available. SB Intuitions

References for Models and Architectures

References for Training Methods

Our Contributors

We love contributors! Feel free to contribute to this project.

contributors

Citation

The summary of this repository is also published as a preprint: Exploring Open Large Language Models for the Japanese Language: A Practical Guide

When referencing this repository, please cite as follows:

@article{awesomeJapanese2024,
    title={{Exploring Open Large Language Models for the Japanese Language: A Practical Guide}},
    author={Kaito Sugimoto},
    doi={10.51094/jxiv.682},
    journal={Jxiv preprint},
    year={2024}
}

Footnotes

  1. Some architectural changes have been made. For details, refer to: 1,000億パラメータ規模の独自LLM「PLaMo-100B」の事前学習

  2. Refer to the following articles: 大規模言語モデルTanuki-8B, 8x8Bの位置づけや開発指針など, 大規模言語モデルを開発するにあたっての事前・事後学習の戦略メモー特に合成データについてー 2

  3. Despite the license, the model card describes the model as a research-purpose release and states that commercial or mission-critical use is not intended.

  4. Some performance enhancements have been made to the original Llama model. See here for details.

  5. Details have not been made public but the private dataset includes data from the EleutherAI Polyglot project's Japanese team and from members of Stable Community Japan.

  6. This project conducted evaluation research on using right-to-left generation instead of the usual left-to-right generation, releasing both left-to-right and right-to-left models.

  7. Despite the license, the model card states that use is intended only for research and development purposes; direct use in actual clinical settings for disease diagnosis or clinical decision-making support is not recommended. 2 3 4 5 6 7 8 9 10

  8. Distributed for research purposes only; redistribution, commercial use, and direct use in actual clinical settings are not permitted. The model is also still at the base (pre-training) stage and has known limitations such as unstable end-of-sequence (EOD / EOS) token output and repetition.

  9. Before conducting Instruction Tuning, a Chat Vector between Llama 3 Instruct and Llama 3 Base is added. 2

  10. After conducting Instruction Tuning, a Chat Vector between Llama 3 Instruct and Llama 3 Base is added. 2

  11. However, if commercial use of KARAKURI LM is desired, direct contact with the developer, KARAKURI Inc., is required.

  12. In Instruction Tuning, because it uses data generated by OpenAI's models, such as GPT-3.5 and GPT-4, for training, there is a possibility that it may violate OpenAI's terms. 2 3 4 5 6 7 8 9 10

  13. Before conducting Instruction Tuning, a Chat Vector between Gemma 2 Instruct and Gemma 2 Base is added.

  14. Despite the license, the model card does not recommend direct use in actual clinical settings for disease diagnosis or clinical decision-making support, and recommends limiting use to an information-providing tool that assists the judgment of medical professionals. 2 3

  15. Although the base model is not officially disclosed, the config.json architecture is DeepseekV3ForCausalLM, the tokenizer is identical to DeepSeek-V3, and the repository includes a DeepSeek NOTICE file, strongly suggesting it is based on DeepSeek-V3.

  16. ○: The model is on the HuggingFace Model Hub and can be loaded in with the AutoModel.from_pretrained() command. △: The model is not on the Model Hub but can be loaded in manually with the HuggingFace transformers library. ✕: The model is not directly loadable with HuggingFace.

  17. By removing Causal Attention from Llama, it is used as an encoder-type model. 2

  18. This project conducted evaluation research on pre-tokenization morphological analysis and released their best performing model, which used Juman++ and BPE.

  19. However, the maximum sequence length has been extended to 2048, and various architectural changes have been made compared to the original BERT. See the HuggingFace repository README for details.

  20. nlp-waseda/roberta-base-japanese and nlp-waseda/roberta-large-japanese trained using a 128 token context length, but nlp-waseda/roberta-large-japanese-seq512 expanded the context length to 512.

  21. Extended to a 1282 context length from the usual 512.

  22. For details of each model, please refer to Chapter 4 of the authors' paper. Note that the SC-2M-wiki model is strictly not a domain-specific model as it is pre-trained only on Wikipedia.

  23. The "small" model trains on Japanese Wikipedia and the Japanese Financial Corpus simultaneously, while the "base" model takes the TohokuUniversityBERT and conducts additional training on the Japanese Financial Corpus.

  24. ManbyoWordPiece conducts a pre-tokenization step using MeCab (IPA+Manbyo dictionaries) and uses WordPiece for subword tokenization, while the SentencePiece model tokenizes text directly using a unigram model.

  25. The classification of embedding models was referenced from Dense Text Retrieval based on Pretrained Language Models: A Survey (Zhao+, 2022). The Bi-Encoder architecture inputs two separate inputs into the model and vectorizes each, using their dot product or cosine similarity as a measure of their proximity. In contrast, the Cross-Encoder architecture inputs the combined inputs into the model to directly compute their proximity internally. Although Cross-Encoders incur higher computational costs, they are often used as rerankers in information extraction due to their ability to compute input proximity more precisely. Among Bi-Encoders, there are types (e.g., ColBERT) that represent the input as multiple vectors (such as one per token) rather than a single vector, hence further classification into Single-representation bi-encoders and Multi-representation bi-encoders.

  26. However, it calls for consideration for use in research and education. Additionally, be aware that some of the licenses for the source models are not Apache 2.0. 2 3

  27. This model transcribes speech into katakana sequences rather than ordinary Japanese orthography. It was developed as the backbone of Kana-CER, a metric for measuring the kanji reading accuracy of TTS systems.