Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis

Lian, Shiguo; Zhao, Kaikai; Lei, Xuejiao; Wang, Ning; Long, Zhenhong; Yang, Peijun; Hua, Minjie; Ma, Chaoyang; Liu, Wen; Wang, Kai; Liu, Zhaoxiang

Computer Science > Artificial Intelligence

arXiv:2502.11164 (cs)

[Submitted on 16 Feb 2025]

Title:Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis

Authors:Shiguo Lian, Kaikai Zhao, Xuejiao Lei, Ning Wang, Zhenhong Long, Peijun Yang, Minjie Hua, Chaoyang Ma, Wen Liu, Kai Wang, Zhaoxiang Liu

View PDF HTML (experimental)

Abstract:DeepSeek-R1, known for its low training cost and exceptional reasoning capabilities, has achieved state-of-the-art performance on various benchmarks. However, detailed evaluations from the perspective of real-world applications are lacking, making it challenging for users to select the most suitable DeepSeek models for their specific needs. To address this gap, we evaluate the DeepSeek-V3, DeepSeek-R1, DeepSeek-R1-Distill-Qwen series, and DeepSeek-R1-Distill-Llama series on A-Eval, an application-driven benchmark. By comparing original instruction-tuned models with their distilled counterparts, we analyze how reasoning enhancements impact performance across diverse practical tasks. Our results show that reasoning-enhanced models, while generally powerful, do not universally outperform across all tasks, with performance gains varying significantly across tasks and models. To further assist users in model selection, we quantify the capability boundary of DeepSeek models through performance tier classifications and intuitive line charts. Specific examples provide actionable insights to help users select and deploy the most cost-effective DeepSeek models, ensuring optimal performance and resource efficiency in real-world applications.

Subjects:	Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2502.11164 [cs.AI]
	(or arXiv:2502.11164v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2502.11164

Submission history

From: Kaikai Zhao [view email]
[v1] Sun, 16 Feb 2025 15:29:58 UTC (7,177 KB)

Computer Science > Artificial Intelligence

Title:Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators