Next Article in Journal
An Enhanced Siamese Network-Based Visual Tracking Algorithm with a Dual Attention Mechanism
Next Article in Special Issue
Exploring Tabu Tenure Policies with Machine Learning
Previous Article in Journal
Federated Learning-Based Location Similarity Model for Location Privacy Preserving Recommendation
Previous Article in Special Issue
Investment Portfolios Optimization with Genetic Algorithm: An Approach Applied to the Spanish Market (IBEX 35)
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Beyond the Benchmark: A Customizable Platform for Real-Time, Preference-Driven LLM Evaluation

by
George Zografos
and
Lefteris Moussiades
*
Department of Informatics, Democritus University of Thrace, 65404 Kavala, Greece
*
Author to whom correspondence should be addressed.
Electronics 2025, 14(13), 2577; https://doi.org/10.3390/electronics14132577
Submission received: 1 May 2025 / Revised: 21 June 2025 / Accepted: 24 June 2025 / Published: 26 June 2025
(This article belongs to the Special Issue Advances in Algorithm Optimization and Computational Intelligence)

Abstract

The rapid progress of Large Language Models (LLMs) has intensified the demand for flexible evaluation frameworks capable of accommodating diverse user needs across a growing variety of applications. While numerous standardized benchmarks exist for evaluating general-purpose LLMs, they remain limited in both scope and adaptability, often failing to capture domain-specific quality criteria. In many specialized domains, suitable benchmarks are lacking, leaving practitioners without systematic tools to assess the suitability of LLMs for their specific tasks. This paper presents LLM PromptScope (LPS), a customizable, real-time evaluation framework that enables users to define qualitative evaluation criteria aligned with their domain-specific needs. LPS integrates a novel LLM-as-a-Judge mechanism that leverages multiple language models as evaluators, minimizing human involvement while incorporating subjective preferences into the evaluation process. We validate the proposed framework through experiments on widely used datasets (MMLU, Math, and HumanEval), comparing conventional benchmark rankings with preference-driven assessments across multiple state-of-the-art LLMs. Statistical analyses demonstrate that user-defined evaluation criteria can significantly impact model rankings, particularly in open-ended tasks where standard benchmarks offer limited guidance. The results highlight LPS’s potential as a practical decision-support tool, particularly valuable in domains lacking mature benchmarks, offering both flexibility and rigor in model selection for real-world deployment.
Keywords: LLM evaluation; user-defined evaluation criteria; customizable benchmarking framework; multi-model comparison; prompt engineering; domain-specific NLP LLM evaluation; user-defined evaluation criteria; customizable benchmarking framework; multi-model comparison; prompt engineering; domain-specific NLP

Share and Cite

MDPI and ACS Style

Zografos, G.; Moussiades, L. Beyond the Benchmark: A Customizable Platform for Real-Time, Preference-Driven LLM Evaluation. Electronics 2025, 14, 2577. https://doi.org/10.3390/electronics14132577

AMA Style

Zografos G, Moussiades L. Beyond the Benchmark: A Customizable Platform for Real-Time, Preference-Driven LLM Evaluation. Electronics. 2025; 14(13):2577. https://doi.org/10.3390/electronics14132577

Chicago/Turabian Style

Zografos, George, and Lefteris Moussiades. 2025. "Beyond the Benchmark: A Customizable Platform for Real-Time, Preference-Driven LLM Evaluation" Electronics 14, no. 13: 2577. https://doi.org/10.3390/electronics14132577

APA Style

Zografos, G., & Moussiades, L. (2025). Beyond the Benchmark: A Customizable Platform for Real-Time, Preference-Driven LLM Evaluation. Electronics, 14(13), 2577. https://doi.org/10.3390/electronics14132577

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop