JOR-Bench: New Japanese-Language Benchmarks Test AI's Operations Research Skills
Researchers introduced JOR-Bench, a collection of five Japanese-language benchmarks with 1,319 problems to evaluate how well large language models (LLMs) can formulate and solve operations research (OR) problems, covering linear programming, mixed-integer programming, non-linear programming, and combinatorial optimization.

Researchers have developed JOR-Bench, a collection of five Japanese-language benchmarks designed to evaluate how well large language models (LLMs) can formulate and solve operations research (OR) problems. Each benchmark is a Japanese translation of an existing English benchmark: IndustryOR, MAMO Complex LP, NL4OPT, OptiBench, and OptMATH. Together, they cover 1,319 problems spanning linear programming, mixed-integer programming, non-linear programming, and combinatorial optimization. JOR-Bench is solver-independent, meaning it can be used with any solver or programming language.
This development matters because it allows businesses and researchers to better understand how AI can be applied to real-world operational challenges in Japanese. For example, companies might use AI to optimize supply chains, manage resources, or solve complex scheduling problems. Having a standardized benchmark in Japanese ensures that AI tools can be reliably tested and improved for practical use.
The benchmarks are based on existing English resources like IndustryOR and OptMATH, which are available online and can give you a sense of the types of problems AI is being tested on. If you're interested in AI's problem-solving capabilities, checking out these benchmarks is a great starting point.