Next Article in Journal
Correction: Chaudhary and DaSouza (2024). Consumer Readiness for Microtransactions in Digital Content Business Models. Businesses, 4(3), 473–490
Previous Article in Journal
Legal Certainty in Digital Commercial Contracting: A Quantitative Contract-Level Study in the Ecuadorian Context
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Benchmarking Generative AI Models for Skill-Aligned Job Posting Generation: A Multi-Domain Comparative Evaluation

by
Alexandros Adam
,
Konstantinos Georgiou
* and
Lefteris Angelis
School of Informatics, Aristotle University of Thessaloniki, 54124 Thessaloniki, Greece
*
Author to whom correspondence should be addressed.
Businesses 2026, 6(3), 43; https://doi.org/10.3390/businesses6030043
Submission received: 15 July 2026 / Revised: 4 August 2026 / Accepted: 6 August 2026 / Published: 14 August 2026

Abstract

Generative Artificial Intelligence (Gen AI) is reshaping recruitment, with organizations increasingly using it to draft job postings and screen candidates. Research on labor market analytics has concentrated on extracting skills from existing postings; the complementary task of generating postings aligned with a required skill set has received little systematic attention. This paper addresses that gap by benchmarking ten Gen AI models on skill-aligned job posting generation across three ESCO-derived skill domains: Finance, Healthcare, and Craft and related trades workers. Each model produced 100 postings per domain, giving 3000 synthetic postings, evaluated with five complementary metrics—Perplexity, Cosine Similarity, Word Mover’s Distance, LIWC-style divergence, and Explicitness—and verified through inferential statistical testing. No single model dominated across metrics or domains. The models prioritized explicit skills over the broader contextual information found in real postings, and performance showed a marked domain sensitivity, with a significant model-by-domain interaction for four of the five metrics. These findings offer guidance for HR practitioners and business decision-makers selecting a Gen AI tool for recruitment tasks. Recruitment outcomes such as time-to-hire and candidate–job fit were not measured, so such benefits remain potential practical implications rather than empirically demonstrated effects.

1. Introduction

In the rapidly evolving contemporary labor market, it is essential to understand the needs and demands of the market for organizations, businesses, and general stakeholders to make informed decisions regarding potential candidates. Inaccurate or poorly written job postings attract a high volume of unqualified applicants, raising screening costs and slowing the hiring process; this inefficiency is a tangible business cost that well-targeted, skill-aligned postings can help reduce (Mahjoub & Kruyen, 2021). An efficient way to recognize the needs of the job market lies in the alignment of skills with job descriptions (Khaouja et al., 2021); this process enhances the accuracy of the reflection of the job requirements with the job role.
Along with the continuous evolution of the labor market, artificial intelligence has also emerged with extraordinary velocity in recent years. Multiple sectors have been influenced by the utilization of AI; it is now among the most influential technologies across research and industry (Eloundou et al., 2024). Among the huge range of applications that AI systems cover, one technology that has made tremendous success in recent years is Gen AI. The term Gen AI refers to a technology that is capable of generating new, meaningful, and rich content such as text, images, or even audio (Cao et al., 2023). This groundbreaking technology has, in the last few years, drastically changed the way we operate in our jobs and how we communicate (Eloundou et al., 2024).
The labor market has evolved unexpectedly in recent years; the workforce and economic development depend on understanding and adapting to these changes. The prevalence of the internet in the last few decades has led to digitalized job postings on various job portals. Analyzing online job postings constitutes a real challenge for many market analysts who attempt to capture patterns or trends in real time. Occasionally, external factors significantly affect the labor market, such as the COVID-19 pandemic, which altered job market demands and the need for employees to acquire skills that were not previously necessary (Forsythe et al., 2020). Therefore, it is imperative to correctly align skills with the current demands of the job market, as they rapidly change over certain time periods.
However, despite the extensive research on extracting skills and analyzing job postings from online sources, there is a notable gap in generating skill-aligned job postings that accurately reflect required skill sets. Recent advances in artificial intelligence are helpful in meeting the challenges that arise due to the rapid evolution of the labor market, and the emergence of new skill requirements. The use of Gen AI models has increased across various domains in recent years in order to solve complex tasks where traditional approaches encounter restrictions; thus, these models are called upon to resolve the existing limitations in skill-related labor market analysis. Consequently, the generative AI approach highlights a new path for research development that will narrow down these limitations in the dynamic environment of the labor market. Hence, the present study is guided by the following four research questions, each of which is addressed empirically in Section 3 and Section 4. The corresponding methodological stages through which each question is answered are described immediately below.
RQ1. 
Can Gen AI models generate job postings that are demonstrably aligned with an explicitly defined set of ESCO domain skills, and how closely do the resulting postings resemble real-world postings from the same domain? As the first step in the research, we introduce a framework for generating job postings produced by an explicitly defined list of domain skills. Unlike other prior studies which primarily focus on extracting and classifying skills from job descriptions, we attempt to broaden labor market analysis by generating occupations that are accurately aligned with skills.
RQ2. 
Does generation quality differ systematically across skill domains and is there a statistically significant model-by-domain interaction? At a second level, multiple skill domains were selected to enhance the comparison clarity and fairness of models in terms of performance, efficiency, and assessment of the model–domain relation according to results. The approach of the current research aims to examine how models adapt across different domains, providing insights about the robustness of the models.
RQ3. 
Under identical prompts and skill sets, do the ten models differ significantly from one another, and does any single model dominate across domains? At a third level, the numerous model selection enables fair comparison and benchmarking among the models, while preserving equality with identical prompts and skill sets as inputs. Existing studies have used a limited number of models, which leads to a lack of comprehensive results and systematic comparisons.
RQ4. 
Are conclusions about model ranking stable across five complementary evaluation metrics and across parametric and non-parametric tests? At the last level, the examination of the outcomes relies on multiple metrics providing a multi-dimensional evaluation and also on statistical analysis to preserve the quality and reliability of the results.
Against this background, the four original contributions of this paper are as follows. They are stated explicitly here because each corresponds to a gap that the studies reviewed in Section 2 overlook.
(i)
A generation framework that conditions job posting generation on an explicitly defined, taxonomy-grounded skill set derived from ESCO, so that the skills a posting should express are fixed in advance and the generated text can be evaluated against them rather than judged in the abstract.
(ii)
The first multi-model benchmark of this task, comparing ten generative models under an identical prompt and an identical skill set. Prior work, including JobSkape, generates postings with a single pipeline and treats generation as a means to another end; here the generated posting is itself the object of measurement, which is what makes a comparison between models possible.
(iii)
A multi-domain design that treats the skill domain as an experimental factor rather than a setting and therefore allows a model-by-domain interaction to be tested directly. This is the source of the study’s central empirical finding, namely that model advantage is domain-dependent rather than global.
(iv)
A combined evaluation protocol that pairs five complementary automatic metrics, spanning linguistic fluency, semantic alignment, stylistic resemblance and skill explicitness, with inferential statistics and effect sizes, so that observed differences between models are shown to be systematic rather than incidental.
Taken together, these constitute a shift from asking whether a generative model can write a plausible job posting to asking which model writes postings that are faithful to a required skill profile, and in which occupational context—a question that has direct bearing on how such tools are selected in practice.
The rest of the paper is organized as follows: In Section 2, we provide some useful background information and fundamentals regarding large language models, applications that LLMs cover, the role of LLMs in labor market analysis, and lastly their utilization in skill extraction. In Section 3, we present the methodological process followed in this research and analyze the steps involved in generating job postings with generative AI models, while Section 4 illustrates the results of experiments and explains the findings. Finally, Section 5 discusses the practical implications and Section 6 includes some closing remarks and conclusions related to the research undertaken.

2. Background and Related Work

This section provides information and fundamental definitions regarding large language models (LLMs) to provide a better understanding of the context of the subsequent research content. The first part contains some useful information and necessary definitions related to LLMs and the applications they cover in the industry, while the second part briefly analyzes the influence of LLMs in labor market analytics, and finally the third part presents the utilization of LLMs regarding skill extraction.
Large language models belong to a category of deep learning models trained on enormous amounts of data, capable of handling complex language tasks such as translation, summarization, information retrieval and much more. LLMs are built on the transformer architecture, a type of neural network architecture that surpasses others in terms of handling sequences of words and capturing patterns in text (Vaswani et al., 2017). In other words, LLMs operate as huge statistical prediction machines which repeatedly predict the next word within a sequence of words. The groundbreaking architecture of these models can produce text that is to human-level performance in a variety of tasks.
Large language models are general-purpose models adept at performing a wide range of natural language-processing tasks. Core capabilities of these models consist of natural language generation, answering questions, text summarization, text classification, language translation, and information extraction. Moreover, virtual assistants such as chatbots have arisen in recent years, which are used in multiple online platforms to provide assistance and useful information to humans instead of a real person. Semantic search is another crucial feature of LLMs with the ability to capture the meaning and underlying intentions of user queries, enabling rapid and accurate responses in information retrieval and search tasks. Text is not the sole method to communicate with LLMs; it is possible via voice and speech recognition, a tool that is present in nearly all smart devices. The user, by utilizing voice, can issue commands to a machine in real time, allowing operations to be completed without the user physically touching a screen (Zhao et al., 2023).
The labor market has experienced a significant transformation due to widespread online job postings. Contemporary job portals offer flexibility in searching for a particular posting by providing multiple features to effectively search for a job according to specific qualifications such as skill levels, occupation categories, and part- or full-time roles, representing a useful information resource for analysts and human resource personnel. Analysts attempting to handle such a large amount of data need tools and applications that can handle extensive data to derive meaningful insights. Large language models specialize in handling enormous amounts of data effectively and can improve labor market analysis by identifying trends and patterns related to the availability of specific skills, wage trends, regional disparities, and diversity matters. A growing body of evidence indicates that organizations are already adopting generative AI tools to automate recruitment-related tasks such as candidate screening and job-description drafting (Abdelhay et al., 2025).
Beyond the technical literature, the business and human-capital management literature explains why this task matters economically. Persistent skills gaps and skill mismatches raise both the cost and the duration of hiring (Cappelli, 2015). Talent-acquisition costs are a recognized operational burden for firms (Tuttle & Critchlow, 2025), and recent evidence indicates that generative AI can raise the productivity of knowledge and support work, including HR-related writing tasks (Brynjolfsson et al., 2025). Situating job posting generation within this economic context clarifies its business value: better-targeted, skill-aligned postings are associated in this literature with lower screening effort, shorter time-to-hire and stronger employer branding. These downstream outcomes are not measured in the present study, which evaluates the linguistic and semantic quality of generated postings rather than realized recruitment performance.
Despite the advancements so far in labor market analysis, the primary objectives of current workarounds mostly focus on extracting and analyzing information from job postings rather than generating job postings that are accurately skill-aligned with the corresponding descriptions. The current study aims to fill this research gap by introducing a complete investigation into the generation of job postings, utilizing LLMs by supplying domain skill sets as inputs.
The job market is constantly evolving at a rapid pace, and this is accompanied by an urgent demand for skills. Analysis of the job market requires useful insights from online job postings, and to accomplish that, it needs a skill extraction process. The term skill extraction refers to the process where skills are gathered from various sources in an unstructured format, such as job postings or resumes, in order to collect the competences that will be used to identify trends or patterns in labor market analysis (Khaouja et al., 2021). The outcome of the analysis with the skill extraction is to provide a general overview of the job market to analysts or stakeholders, and the areas where they should mainly focus and apply their resources to be competitive in the market.
In an attempt to eliminate the restrictions of the traditional approaches, the LLM-based models were introduced to address the issues of earlier methods. The LLM-based methods are capable of learning from large-scale corpora to distinguish explicit and implicit skill requirements that are included in job descriptions. The great advantage of LLMs is capturing the semantic relationship among skills, occupations, and operations to provide a more accurate and effective skill extraction process. Moreover, the dependency on labeled datasets is reduced by applying zero-shot and few-shot techniques across various domains when extracting skills with limited human supervision (Brown et al., 2020). The advent of LLMs in the skill extraction process has helped to address and overcome traditional limitations and adapt to the dynamic environment of the labor market and its demand for skills.
The closest prior work to the present study is JobSkape (Magron et al., 2024), a framework that generates synthetic job postings to support skill-to-taxonomy matching. The present study differs from JobSkape in four concrete respects. First, JobSkape generates synthetic postings primarily as training material to improve downstream skill-to-taxonomy matching, whereas here the generated posting is itself the object of evaluation. Second, JobSkape is not designed as a cross-model comparison, whereas this study benchmarks ten generative models under identical prompts and skill sets. Also, this study evaluates three deliberately heterogeneous skill domains and tests explicitly for a model-by-domain interaction, so that domain sensitivity is quantified rather than assumed. Finally, instead of a single matching-oriented criterion, five complementary metrics spanning linguistic fluency, semantic alignment, stylistic resemblance and skill explicitness are applied, and the resulting differences are validated with inferential statistics. The contribution of this paper is therefore a multi-model, multi-domain and multi-metric benchmark of skill-conditioned job posting generation, rather than a generation pipeline for augmenting skill-matching datasets.

3. Methodology

In this section, the applied research methodology pipeline is presented in detail to provide a clear and structured overview of each step of the process. Starting with the first part, it illustrates the end-to-end process through a flow diagram. The next part describes the whole process of skill extraction and the creation of domain skill sets, while the third part contains information about the ground truth dataset construction. The fourth part demonstrates the process of job posting generation utilizing generative AI models, including the prompt designs that were inserted as inputs. Part five analyzes the evaluation metrics employed for this research, while part six contains the statistical analysis used to validate the experimental results. The last part includes a qualitative comparison between synthetically generated and real-world job descriptions.
The proposed methodology of the current research is structured as a multi-stage pipeline and organized into well-defined steps. These steps include: (i) skill extraction and the creation of skill domain sets, (ii) the generation of synthetic job postings, (iii) the construction of ground truth datasets, (iv) quantitative evaluation of the metric results, (v) verification of the metric results via statistical analysis and (vi) qualitative comparison between synthetic and real-world job postings. Figure 1 illustrates the complete pipeline flow of the methodology, where each step of the process cycle is clearly visible. The pipeline cycle begins with the extraction of skills from the ESCO taxonomy, which is the main resource for collecting skills across three different domains. The extracted skills are applied for the creation of three distinct sets of skills according to a specific domain area. These sets of skills are used to design descriptive and well-instructed prompts for each domain that act as inputs into the generative AI models.
Following the design of the prompts, they are applied to ten different generative AI models to produce synthetic job postings according to the specific requirements and characteristics of each domain. Each model generated 100 job postings per domain, giving a total of 3000 synthetic postings across the ten models and three domains; this per-cell sample size underlies every statistical test reported in Section 4. In the interim, the construction of the ground truth datasets takes place and consists of real-world job postings corresponding to selected domains. The ground truth datasets serve as a benchmark for evaluation metrics and for visual comparison with the produced content. In the final phases of the methodology, the primary focus is on evaluating the quality of the results. This is initially achieved by performing a quantitative evaluation with five different metrics. In addition, a statistical analysis is conducted to verify the reliability of the evaluation and ensure that the findings of the research are statistically correct, and finally a qualitative comparison of generated job postings with real-world examples is performed with the intention of identifying any concealed remarks that evaluation metrics are unable to detect.
The primary source for skill extraction in this research is ESCO (European Skills, Competences, Qualifications and Occupations) which operates as a standardized taxonomy for describing, identifying and classifying professional occupations and skills in the European labor market (le Vrang et al., 2014). The specific ESCO release used, the criteria applied when placing skills into each domain skill set, and the procedures used to remove duplicate and cross-domain overlapping skills are reported in Appendix A.2. The ESCO platform is widely used for matching jobseekers based on their skills, as well as suggesting training content for individuals who seek to reshape an existing skill or acquire a new one. Following the selection of the skill source, the next stage of the process relies on the selection of domain areas. The primary criterion of selection is diversity, ensuring that chosen domains are substantially different in terms of background knowledge, terminology and required competencies. The choice path aims to prevent skill overlap among domains and enable robust evaluation of Gen AI models across divergent skill environments. Based on these criteria, three distinct domains were chosen: Finance, Healthcare and Craft and related trades workers. These three domains are not defined at an identical level of taxonomic granularity: “Craft and related trades workers” corresponds to a major ISCO-08 group, whereas Finance and Healthcare were operationalized as broader sectoral groupings of ESCO occupations. This asymmetry follows from the decision to contrast heterogeneous linguistic environments rather than taxonomically equivalent units, but it is a genuine limitation of the domain comparison and is treated as such in Section 5 and Section 6. The taxonomic definition, the number of occupations and skills, and the number of reference postings for each domain are reported in Appendix A.2.
The three domains were not selected as a representative sample of the labor market, which no set of three domains could be, but as a deliberately contrasting set chosen along the dimensions expected to affect generation. Finance and Healthcare are both linguistically formal and terminologically dense, but they differ in the kind of density involved: Finance vocabulary is heavily procedural and regulatory, whereas Healthcare vocabulary is clinical and, in several European systems, tied to protected professional titles. Craft and related trades workers was selected as the deliberate counterpoint, since its postings describe concrete, observable tasks in comparatively standardized language. The three domains also differ in labor-market structure, spanning predominantly white-collar, regulated professions and manual occupations. This spread was chosen so that any domain effect observed would more likely reflect genuine linguistic and structural differences than incidental variation between neighboring sectors. It follows that the results should be read as evidence that generation quality is domain-sensitive, and as an indication of the direction that sensitivity takes, rather than as coefficients transferable to any particular unexamined domain. Whether the ordering observed here holds for domains with different characteristics—for example creative, legal or educational occupations, or occupations whose postings are typically short—is an empirical question this design cannot answer, and extending the benchmark across a wider range of occupational fields is identified as a direction for future work in Section 6.
The real occupation postings were retrieved via an application programming interface (SKILLAB API) provided by the university for research purposes, serving as a genuine resource for labor market data. Once the data retrieval was complete, a sequence of preprocessing steps were performed to construct the final ground truth datasets. Initially, missing values in the skills column were identified and removed in order to keep the consistency within the dataset. Subsequently, the next step was to transform Uniform Resource Identifiers (URIs) into human-readable text. This transformation was achieved by applying two dictionaries as references derived from the ESCO dataset. The first dictionary was created to map the skill URIs to skill names while the other mapped occupation URIs to occupation titles. Following the URI-to-text conversion, unnecessary attributes of the dataset were discarded, leaving only the essential fields for analysis such as job title, job description and associated skills. In addition, a text normalization step was applied, including the conversion of the text content to lowercase and the removal of HTML tags from the job description attribute field. In the final step of the process, the cleaned dataset was filtered based on the domain criteria, resulting in the generation of three unique ground truth datasets corresponding to the chosen skill domains. Throughout this paper the term “ground truth” is used in its machine-learning sense, denoting a fixed empirical reference against which generated text is compared; it does not imply that these postings are optimally written, and the terms “ground truth dataset” and “reference dataset” are used interchangeably. The composition of the three reference datasets is reported in Table A3.
A note on the source of the reference data is warranted. The reference postings used in this study were obtained through the SKILLAB API, which aggregates online vacancy advertisements published across European labor markets, and not from LinkedIn or comparable commercial recruitment platforms. Three considerations led to this choice. First, the terms of service of such platforms do not permit the systematic collection and redistribution of posting content for research purposes, and no comparable public research license was available to the authors. Second, the reference corpus had to be linkable to ESCO occupations and skills for the skill-conditioned design used here, which the SKILLAB pipeline supports directly. Third, a platform-independent source avoids conflating model behavior with the editorial and formatting conventions of a single commercial platform. A comparison against postings drawn from a major recruitment platform would nevertheless be a natural and valuable extension of this work, since platform-specific conventions may themselves affect what counts as a well-formed posting, and it is identified as future work in Section 6.
Following the construction of ground truth datasets is the design of prompts and the selection of generative AI models. The prompt design was constructed to ensure the quality of generated job postings in terms of structure, descriptive detail and skill alignment with selected domains. Three domain-specific prompts were developed for the current research, corresponding to the previously selected skill domains. All prompts follow a unified design framework to preserve consistency across experiments. Each prompt begins with a role-based instruction that assigns the task of acting as a creative copy-writer with expertise in the specific domain. This role framing was adopted because persona-style instructions stabilize register and output structure in instruction-tuned models, and because the target artifact is a public-facing recruitment advertisement whose register is closer to marketing copy than to technical documentation. An identical role instruction was used for every model and every domain, so that whatever effect it has is held constant across the comparison. The indicative prompts for the three domains, including the few-shot examples and skill lists actually used, are provided in Supplementary Material S1 (Tables S1.1–S1.3). This role definition establishes the expected writing tone and level of professionalism in the generated results. Furthermore, detailed instructions were provided in bullet-point format to specify the mandatory components of each job posting. These components included a job title, a comprehensive role description, and a list of required skills. Additional constraints instructed the models to utilize at least two skills from the predefined skill lists created in earlier stages of the methodology. Finally, formatting instructions were included to ensure that the generated outputs followed a structured CSV format, facilitating efficient handling of results in subsequent stages of the methodology. To enhance the quality of the generated text and minimize ambiguity, each prompt included three real-world job posting examples related to the corresponding skill domain. These examples served as few-shot demonstrations to guide the models with respect to content structure, domain-specific terminology, and descriptive depth. The three real postings used as few-shot examples in each prompt were drawn from a pool that is separate from the ground-truth set used for the WMD and Cosine Similarity metrics, so the few-shot examples do not contaminate the similarity evaluation. This separation matters: reusing ground-truth postings as few-shot exemplars would have allowed the models to overfit to the very texts against which they were later scored, artificially inflating the similarity metrics and undermining the fairness of the comparison.
A total of ten distinct models were selected for the purposes of the research, enabling a broad and exploratory comparison of job posting generation across diverse model architectures, providers, and deployment environments. The model selection was not based on strict performance criteria but rather aimed to include a diverse set of well-known generative systems, incorporating proprietary as well as open-source solutions. The selection was designed to span a diverse range of architectures (dense transformer decoders as well as mixture-of-experts designs), parameter scales, and development origins (proprietary versus open-weight); this diversity was sought deliberately in order to examine the trade-off between model complexity, licensing cost, and output quality—a factor of direct relevance to business adoption. This strategy allows the research to study variability in model behavior under uniform prompt conditions during evaluation. The chosen models were accessed via diverse deployment mechanisms, incorporating cloud-based application programming interfaces (APIs) and local environment setup. The selected models are: Gemini-2.5-flash developed by Google (Mountain View, CA, USA); a Mistral instruction-tuned model developed by the French company Mistral AI (Paris, France) (the exact checkpoint is reported in Table S2.1); Grok 4.1 developed by xAI (Palo Alto, CA, USA); NVIDIA: Nemotron Nano developed by NVIDIA (Santa Clara, CA, USA); Hermes 3 by Nous Research (New York City, NY, USA); GPT-OSS, OpenAI’s open-weight release (the exact checkpoint and parameter size are reported in Table S2.1); Deepseek-r1 developed by the Chinese startup DeepSeek (Hangzhou, China); Gemma3 developed by Google DeepMind (London, UK); Llama-3.1 developed by Meta (Menlo Park, CA, USA); and Qwen-2.5 developed by Alibaba’s DAMO Academy (Hangzhou, China). To support independent reproduction, the exact identifier, version and access route of every model are reported in Table S2.1, and the decoding configuration in Table S2.2. All models were queried with a single fixed prompt per domain, without retries, reranking or manual post-editing of the returned text, and outputs were parsed programmatically from the requested CSV format.
To evaluate the quality of the synthetically generated job postings, multiple complementary metrics were employed. The selected metrics capture different aspects of text quality, such as linguistic fluency, semantic alignment, stylistic consistency, and skill explicitness. At the quantitative evaluation phase, metric scores were computed at the individual job posting level and then aggregated by calculating the mean values per skill domain and per model. These mean scores were used to construct a metrics table that provides an initial assessment of model performance based on diverse skill domains, while the complete set of individual metric scores was preserved for the statistical analysis in the subsequent stage of the methodology. Five evaluation metrics were adopted in this research; these are: Perplexity, Cosine Similarity, Explicitness, Word Mover’s Distance (WMD) and Linguistic Inquiry and Word Count (LIWC). The Perplexity metric quantifies the level of “uncertainty” a model faces when predicting the next token in a sequence, where a lower Perplexity score indicates that the evaluated text is more predictable under the external reference model. It should be stressed that Perplexity, as used here, is a property of the generated text as judged by that reference model, and not a measure of the confidence of the model that produced the text (Jelinek et al., 1977). Perplexity was computed with a single fixed external reference model applied to every generated posting, so that the scores are comparable across models and no model scored its own output. The reference model, its tokenizer and the preprocessing applied before scoring are specified in Table S2.3. Cosine Similarity measures the degree of similarity between textual data in high-dimensional embedding space, where a higher score value indicates great semantic alignment between textual representations (Reimers & Gurevych, 2019). For this metric each generated posting was embedded and compared with the embedded representation of its associated skill set, and the per-posting similarities were averaged to obtain the domain-level means reported in Table 1. The embedding model is specified in Table S2.3. Explicitness counts how frequently the predefined skill entities of a domain appear verbatim in a generated posting, as formalized in Equation (5) below. Unlike the other four metrics, Explicitness is not monotonically directional: a very low value indicates that the requested skills are barely stated at all, while a very high value indicates repetitive verbatim enumeration of skill labels. It is therefore best interpreted as a descriptive measure of enumeration intensity, for which intermediate values are preferable to either extreme; the downward arrow attached to this metric in Table 1 and Table 2 should accordingly be read as the direction away from repetitive keyword listing rather than as an unqualified criterion of quality. Word Mover’s Distance (WMD) measures the dissimilarity between two documents by calculating the cumulative “travel cost” required to align the words of one document with another in an embedded vector space, where a lower score for the Word Mover’s Distance indicates that the generated job description is very similar semantically to a real job description (Kusner et al., 2015). WMD was computed between each generated posting and the reference postings of the same domain using the word representations specified in Table S2.3, and the resulting distances were averaged per model and domain. The style divergence metric, referred to throughout this paper as the LIWC-style divergence, measures the linguistic divergence between two documents in terms of their tone, style, and psychologically relevant categories, where a lower value for the LIWC-style divergence score indicates that two documents are highly similar in terms of their sentiment and linguistic tone (Tausczik & Pennebaker, 2010). It should be stated explicitly that the proprietary LIWC software and its licensed dictionary were not used: the score reported here was produced by a custom implementation of the same category-divergence principle, as recorded in Table S2.3. The metric is therefore named the LIWC-style divergence throughout this paper, denoting a category-proportion comparison built on the same principle, and the citation identifies the conceptual basis of the measure rather than the instrument applied. The LIWC-style divergence score reported here is computed over the category set specified in Table S2.3 by comparing the category-proportion profile of each generated posting with the mean profile of the reference postings of the same domain.
For precision, the five metrics are defined formally below. Let p denote a generated job posting consisting of tokens x 1     x n , let P e denote the set of postings generated by a given model for domain d , and let S e denote the skill set of domain d .
P P L ( p ) = e ( 1 n ) · i = 1 n log P r e f ( x i | x 1 x i 1 )
c o s ( u , v )   =   ( u   ·   v ) ( u v )
W M D ( p ,   r )   =   min T Σ i j   T i j   ·   c ( i , j ) ,   s u b j e c t   t o   Σ j   T i j   = d i   a n d   Σ i T i j =   d j
L I W C d i v ( p )   =   ( 1 / C | )   ·   Σ c C | f c ( p )     f c ( R e )   |
E ( p )   =   |   {   s     S e   :   s   o c c u r s   v e r b a t i m   i n   p   }   |
In Equation (1), P r e f is the external reference language model of Table S2.3, so that Perplexity measures the predictability of the generated text under a model that did not produce it. In Equation (2), u and v are the embeddings of the generated posting and of its associated skill set respectively, and the per-posting values are averaged over P e to give the domain means of Table 1. In Equation (3), c ( i , j ) is the Euclidean distance between the embeddings of words i and j , T is the transport matrix, and d i and d j are the normalized bag-of-words weights of the generated posting p and of the reference posting r . In Equation (4), C is the set of LIWC categories used, f c ( p ) is the proportion of tokens of p falling in category c , and f c ( R e ) is the mean proportion for that category over the reference postings R e of the same domain. In Equation (5), E ( p ) is the number of distinct domain skill labels appearing verbatim in the posting, and the reported Explicitness value is the mean of E ( p ) over P e .
The statistical analysis process is designed to evaluate the effect of model choice and skill domain on each evaluation metric. Statistical tests are performed independently for every combination of model, domain, and metric in order to determine whether performance differences are statistically significant. The analysis examines three types of effects: the main effect of the model, the main effect of the domain, and the model–domain interaction effect. For each evaluation metric, the analysis pipeline starts with a two-way analysis of variance (ANOVA). This statistical test is applied in order to capture the existence of statistically significant differences across models, domains, and their interaction. In this formulation, the dependent variable is the evaluation metric, while the independent variables are model factor (with 10 levels), domain factor (with 3 levels) and the interaction effect of the model-domain. The homogeneity-of-variance assumption was violated for every metric. Because the two-way ANOVA F-test assumes homoscedasticity, the F-statistics reported in this study are interpreted as descriptive indicators of the relative magnitude of the model, domain and interaction effects rather than as exact inferential tests, and the significance levels attached to them are read with that caveat. Two safeguards are applied in compensation. First, a rank-based Kruskal–Wallis test, which requires neither normality nor homogeneity of variance, is used as an independent check on the model factor. Second, an effect size is reported for every effect, so that the magnitude of each effect can be judged separately from its p-value: partial eta-squared for the ANOVA terms, computed as η 2 p   =   ( F   ×   d f e f f e c t )   /   ( F   ×   d f e f f e c t   +   d f e r r o r ) , and epsilon-squared for the Kruskal–Wallis terms, computed as ε 2   =   ( H     k   +   1 )   /   ( N     k ) with k groups and N observations. These values are reported in Table 3. No heteroscedasticity-robust two-factor test was carried out, and this is acknowledged as a limitation in Section 6. Because the Perplexity distribution is strongly right-skewed, the divergence between the parametric and the rank-based result for this metric is examined separately in Section 4. Welch’s ANOVA was not adopted because it does not extend cleanly to a two-factor design. The last step of the process includes the examination of the results to identify which models performed better overall and which domain facilitates better performance across models.
The final stage of the methodology involves qualitative comparison among synthetically generated and real-world job descriptions. The purpose of this final step is to examine the output produced by models in contrast with real-world job postings to identify differences that quantitative evaluation metrics may fail to capture. By incorporating a qualitative analysis, the methodology provides a comprehensive perspective to the statistical evaluation, enhancing a more complete understanding regarding the results of the best-performing models.

4. Results

This section focuses on the experimental results of the research and discusses the outcomes derived from the proposed methodological framework. The first phase of presenting the results includes the outcomes of the models in terms of performance in generating job postings across the selected domain using the mean values of each evaluation metric for every combination of model–domain. Subsequently, the results of the statistical analysis are reported as produced by the full evaluation metric values and not solely the mean, providing the significance differences in performance between the models and finally the qualitative analysis of best-performing models with a real-world job posting.
Table 1 and Table 2 report the evaluation results: Table 1 gives the mean value of every metric for each model–domain combination, with rows grouped by model, while Table 2 summarizes which model attains the best value for each metric within each domain. For clarity purposes, the arrows that appear next to the evaluation metrics’ names indicate the desired optimization direction. For the four directional metrics, an upward arrow (↑) symbolizes higher values corresponding to better performance, as with the Cosine Similarity metric, while a downward arrow (↓) symbolizes lower values corresponding to better performance, as with the Perplexity metric.
The analysis starts with the Perplexity evaluation metric, which provides the linguistic fluency of the generated job postings. Out of all evaluated models, the NVIDIA: Nemotron nano model achieved the lowest Perplexity score, with a mean value of 34.07 in the Craft-related domain, indicating the confidence of the model to generate the highest level of fluent content. The analysis continues with the Linguistic Inquiry and Word Count (LIWC)-style divergence metric, which evaluates stylistic comparison between generated job postings and real-world descriptions. Among all evaluated models, Mistral accomplished the lowest score value of 0.26 in the Finance domain, indicating the model resembles real-world cases in terms of its writing style and linguistic tone.
In terms of semantic alignment, the Cosine Similarity metric measures whether the generated content reflects the contextual meaning of the associated skills. Among all evaluated models, Hermes-3 achieved the highest score value of 0.590 in the Craft-related domain, indicating that the model effectively captured the semantic relationships of generated text with corresponding skills. The Word Mover’s Distance evaluation metric measures the dissimilarity between synthetically generated job descriptions with the real-world descriptions derived from the ground truth dataset. Lower values indicate excellent semantic similarity among the compared documents. Across all model–domain combinations the lowest (best) WMD value of 1.10 was obtained by NVIDIA: Nemotron nano in the Craft-related domain; within the Healthcare domain specifically the lowest value, 1.12, was obtained by Llama-3.1. The Explicitness evaluation metric calculates the frequency of predefined skill entities explicitly appearing within a generated job description. Lower score values for Explicitness reveal better model performance where skills are integrated more implicitly within the text rather than direct or repetitive enumeration. Among all the models, Llama-3.1 achieved the lowest value of 5.75 in the Craft-related domain. As set out in Section 3, however, a very low Explicitness value indicates that the requested skills are barely stated explicitly, so this figure should be read as the lowest enumeration intensity rather than as the highest posting quality. Finally, an overall analysis in the metrics table with mean values illustrates that the generative models accomplished stronger performance in the context related to the Craft-related domain. One plausible explanation is the lower linguistic complexity and more standardized vocabulary of craft and trade occupations. This interpretation is not tested directly here, however, and the Craft advantage may equally reflect the coarser taxonomic granularity of that domain, differences in the volume and quality of the available reference postings, or the skewness of individual metrics; these competing explanations are examined in Section 5 and Section 6. The specialized terminology and domain-specific information required in Finance and Healthcare domains appear to challenge the models in generating content based on these scientific areas, a difficulty that is reflected in the experimental outcomes.
Following the representation of experimental results derived from the evaluation metrics table based on mean values, this part of the research presents the findings of the statistical analysis process. The purpose of this phase of the research is to deliver a more detailed and rigorous examination to identify whether the observed performance patterns are statistically accurate or potentially misleading due to the descriptive nature of outcomes.
Table 3 reports, for each evaluation metric, the two-way ANOVA results for the model and domain main effects and their interaction, together with a Kruskal–Wallis test on the model factor used as a non-parametric robustness check (Kruskal & Wallis, 1952). For four of the five metrics the interaction between the model and the domain is statistically significant, indicating that the results produced by a Gen AI model vary depending on the selected domain; Perplexity is a notable exception and is discussed separately below.
Figure 2 illustrates the overall results produced by the statistical analysis pipeline. Beginning with the Craft-related domain, the model NVIDIA: Nemotron nano achieved the best performance in two out of five evaluation metrics (Perplexity and WMD). In the LIWC-style divergence, Gemini-2.5-flash achieved the lowest (best) value. In the Cosine Similarity metric, the Hermes-3 achieved the best results, while Llama-3.1 recorded the lowest Explicitness value; as Explicitness is not directional, this is reported as the lowest enumeration intensity and not as the best result. In the Financial domain, NVIDIA: Nemotron nano again performed strongly, leading on Cosine Similarity among the four directional metrics, and additionally recording the lowest Explicitness value. However, the best-performing model in the Perplexity and WMD metrics was Gemini-2.5-flash, while Mistral achieved the lowest (best) LIWC-style divergence value in this domain. In the Healthcare domain, Gemini-2.5-flash emerged as the best-performing model, achieving the top score in both Cosine Similarity and the LIWC-style divergence. Llama-3.1 led in WMD, and Qwen-2.5 led on Perplexity; Deepseek-r1 recorded the lowest Explicitness value in this domain, which is reported descriptively rather than as a further metric win.
For the LIWC-style divergence, Cosine Similarity, WMD, and Explicitness, the model and domain main effects and their interaction are all statistically significant under the two-way ANOVA (all p < 0.001), and the Kruskal–Wallis test independently confirms a significant model effect for every one of these four metrics (H = 505.07–1087.62, all p < 0.001). Because the Kruskal–Wallis test is rank-based and does not rely on the normality or equal-variance assumptions that the ANOVA requires, this agreement between two different families of test increases confidence that the observed model and domain effects for these four metrics reflect genuine differences rather than artifacts of a single statistical procedure. The model, domain, and interaction terms carry 9, 2, and 18 degrees of freedom respectively, following directly from the ten model levels and three domain levels used in the design; the associated residual degrees of freedom depend on the per-cell sample size. With n = 100 postings in each model×domain cell (N = 3000), the residual degrees of freedom are 2970.
The two-way ANOVA finds no statistically significant effect of model (F = 0.99, p = 0.44), domain (F = 2.24, p = 0.11), or their interaction (F = 1.30, p = 0.18) on Perplexity, yet the Kruskal–Wallis test on the same data finds a highly significant model effect (H = 1670.39, p < 0.001). This divergence is consistent with the extreme right-skew already visible in Table 1, where Perplexity ranges from 34 (NVIDIA: Nemotron nano, Craft) to 2680.75 (Llama-3.1, Craft): a small number of very large outliers inflate the within-group variance that the ANOVA F-test relies on, to the point where a real group difference is no longer detected, whereas the Kruskal–Wallis test operates on ranks and is far less sensitive to such outliers. The ANOVA result for Perplexity should therefore not be read as evidence that the models perform equivalently on this metric; rather, it demonstrates why a rank-based test is a necessary safeguard for this specific metric, and why the Kruskal–Wallis result is the more trustworthy indicator of a real model effect here. The effect sizes in Table 3 make the same point quantitatively: the ANOVA effects on Perplexity are negligible in magnitude (partial η 2 p ≤ 0.008), whereas the rank-based effect size for the model factor is substantial ( ε 2 = 0.556). A log-transformed re-analysis of Perplexity was not performed and is left to future work. It is also worth noting that, because the Kruskal–Wallis test is a one-way procedure, it was applied only to the model factor in Table 3 and does not provide an equivalent non-parametric check on the domain main effect or the model × domain interaction; the same skewness caveat therefore applies with the same force to the domain and interaction rows for Perplexity, and results for those two terms should be treated with caution until a rank-based or log-transformed re-analysis is available.
An overall analysis of the results reveals that the model Gemini-2.5-flash achieved the highest number of top metric scores across all domains, leading in five of the fifteen model-domain-metric cases, compared to NVIDIA: Nemotron nano which led in four. No single model led in a majority of cases, and the margin between the top two models is narrow; this is discussed further, together with the statistical evidence, in Section 5.
This section provides a qualitative comparison among real-world job postings, used as ground truth data, and the synthetically generated job descriptions produced by the best-performing models. For each selected skill domain, comparative tables are presented that include one representative of a ground-truth job description alongside the outputs generated by the two generative models which were distinguished for their robust performance, namely Llama-3.1 and NVIDIA: Nemotron nano. Placing real and generated job descriptions side by side enables a more detailed examination of linguistic quality, semantic alignment, and human-perceived resemblance, complementing the evaluation metrics presented in previous sections of this chapter. This comparison is explicitly illustrative rather than a systematic human evaluation. For each domain, one reference posting and the outputs of the two models leading on the largest number of metrics were selected, and no independent raters, scoring rubric or inter-rater agreement statistics were involved. The observations below should therefore be read as hypotheses generated by qualitative inspection rather than as measured differences in perceived quality. No structured human evaluation was conducted as part of this study, and none is claimed: the comparison is offered solely as a qualitative illustration of the differences that the automatic metrics cannot capture. A formal evaluation with independent raters, an explicit scoring rubric and an inter-rater agreement statistic is identified as the natural next step for this line of work in Section 6. The selected postings were chosen with a fixed seed and by inspection.
The job description generated in the Craft-related domain and particularly the one produced by NVIDIA: Nemotron nano in Table 4 completely differs from the previous domain comparisons. In contrast to the descriptions generated in the Healthcare and Finance domains, this description includes detailed information regarding the job role and explicitly references the hiring company. The output content is noticeably richer and more extensive than the other generated examples and superficially resembles a human-written posting in structure and register. The produced posting contains well-structured and coherent sentences and would plausibly serve as a first draft for a standard job advertisement in this domain, subject to human review; publication readiness as such was not assessed. The qualitative observations meet the quantitative results demonstrated earlier, where models achieved superior results in terms of performance in the Craft-related domain across multiple evaluation metrics. The better results in the Craft-related domain may be related to the lower linguistic complexity and more straightforward terminology of craft and trade occupations, although, as discussed in Section 5 and Section 6, domain granularity and differences in the reference data provide competing explanations that the present design cannot rule out.
The results indicate that none of the models consistently outperformed all others across all metrics. NVIDIA: Nemotron nano and Gemini-2.5-flash were the two most consistent performers overall (Section 4, Table 1), and the illustrative comparison suggests that NVIDIA: Nemotron nano produced the most plausible posting among the examples inspected in the Craft-related domain.
However, the strong performance of the models is followed by trade-offs across evaluation dimensions, as is illustrated in Figure 3. Specifically, Llama-3.1 provided weaker performance in Perplexity and Cosine Similarity, suggesting limitations in producing linguistically fluent content and preserving strong semantic alignment with real-world job postings. In addition, NVIDIA: Nemotron nano showed comparatively lower performance in the LIWC-style divergence and WMD, revealing a struggle of the model to generate text that is not aligned with a human writing style and lexical semantic similarity. These pinpointed findings emphasize that model strengths are metric-dependent, indicating the necessity to utilize multi-dimensional evaluation to preserve the quality and reliability of synthetically generated job postings. It should also be noted that the Explicitness of NVIDIA: Nemotron nano varies sharply across domains—from 7.00 in Finance (its best) to 29.49 in Craft (the worst value in that domain)—which sits awkwardly with both the stated principle that lower Explicitness is better and the model’s status as the overall winner. This large within-model swing warrants investigation to determine whether it reflects a real effect or a measurement artifact.
Figure 4 illustrates the effect of each domain on the quality of synthetically generated job postings, as computed using the combination of Cosine Similarity and WMD evaluation metrics. The results indicate that the Craft-related domain achieved greater performance across models with the combined value of metrics approaching 0.51, suggesting strong semantic alignment of the generated description with associated skills, and lexical proximity between generated and real-world occupational descriptions. The concrete nature and relatively standardized vocabulary of the Craft-related domain enables model efficiency in generating job postings that capture both semantic meaning and vocabulary usage.
Following the Craft-related domain, the Finance domain showed moderately strong performance with a combined score value of approximately 0.48. This result may be interpreted due to the structured patterns commonly found in job postings related to finance, which facilitates cleaner contextual elements for generative models. In contrast, the Healthcare domain was revealed as the most challenging for job posting generation. The complexity of clinical terminology accompanied by domain-specific precision tone causes the models to be less effective and consistent in such domain areas.
In conclusion, among the selected generative models, Gemini-2.5-flash and NVIDIA: Nemotron nano achieved the strongest and most comparable overall performance, leading in five and four metric–domain combinations respectively. Domain characteristics played an important role in generating quality content, with the Craft-related domain consistently achieving better results across all models, likely due to the lower linguistic complexity and more standardized terminology in contrast to Finance and Healthcare domains. Finally, the illustrative qualitative inspection indicated that, among the examples examined, the NVIDIA: Nemotron nano output in the Craft-related domain was the most plausible relative to a genuine job posting.

5. Discussion

The results in Section 4 show that no single Gen AI model dominates across metrics, domains, or evaluation styles. Gemini-2.5-flash led in five of the fifteen model–domain–metric combinations and NVIDIA: Nemotron nano led in four, a margin of one that is well within the range one would expect from ordinary sampling variation rather than a decisive performance gap. This experience is itself a methodological finding worth stating plainly: declaring an “overall best” model by tallying metric-wins is fragile, sensitive to small errors in how winners are computed, and, even when computed correctly, a five-versus-four margin does not constitute strong evidence of superiority without an accompanying measure of uncertainty.
The statistical analysis in Table 3 reinforces this picture rather than resolving it. For the LIWC-style divergence, Cosine Similarity, WMD, and Explicitness, the two-way ANOVA indicates significant model, domain and interaction effects, and the Kruskal–Wallis test independently corroborates the model effect; because Kruskal–Wallis is a one-way procedure, it provides no non-parametric confirmation of the domain effect or of the model × domain interaction, which therefore rests on the ANOVA alone. Taken together, these results support the qualitative observation that model strengths are metric-dependent rather than uniform. The Perplexity result is the most informative single finding of the statistical analysis: the standard ANOVA fails to detect a model effect that a non-parametric, outlier-robust test detects clearly. Taken together with the domain granularity asymmetry noted in the Limitations—Craft is an ISCO major group while Finance and Healthcare are broad sectors—this suggests that some of the apparent “Craft is easier” effect may be inflated by measurement artifacts (skewed metrics, uneven domain granularity) rather than reflecting a purely linguistic phenomenon. Substantively, Craft-and-trades postings tend to describe concrete, observable tasks in standardized vocabulary, whereas Finance and Healthcare postings rely on abstract, regulated, and highly specialized language; current models reproduce the former far more reliably than the latter, which offers a linguistic explanation for the domain gap that complements the measurement-artifact account above. Effect sizes are reported alongside significance levels in Table 3. Future replications of this study should log-transform Perplexity before analysis and should re-examine the domain comparison with domains matched more closely in granularity; the present domain comparison remains limited by the unequal granularity of the three domains.

5.1. Theoretical and Practical Implications

For labor market analytics research, these findings imply that benchmarking generative models for job posting generation should not rely on single-metric or vote-counting comparisons. A multi-dimensional evaluation protocol, of the kind adopted here, is necessary precisely because linguistic fluency (Perplexity), semantic alignment (Cosine Similarity, WMD), stylistic resemblance (LIWC-style divergence), and skill Explicitness capture different and sometimes conflicting aspects of quality; a model that performs well on one dimension may perform poorly on another, as the trade-off analysis in Figure 3 shows directly.
For practitioners considering Gen AI for drafting job postings, the practical implication is that model choice should be driven by which quality dimension matters most for the intended use case, rather than by a single published “best model” claim, and should be re-validated per domain: the model that performs best in Craft-related postings is not necessarily the best choice for Finance or Healthcare postings, where specialized terminology appears to be harder for current models to reproduce faithfully. Given the qualitative finding that generated postings, particularly from smaller open-weight models, tend to be shorter and less contextually rich than real postings, organizations should treat generative output as a drafting aid that a human reviewer edits and extends, rather than as a publication-ready replacement for a professionally written posting.

5.2. Managerial Implications

For HR managers, these results argue against a one-size-fits-all adoption of Gen AI for drafting job postings. The evidence suggests that cheaper, open-weight models such as NVIDIA: Nemotron nano can already be highly effective for blue-collar and technical roles, where tasks are concrete and vocabulary is relatively standardized. For white-collar professional roles in Finance and Healthcare, whose language is more abstract and laden with regulatory and domain-specific terminology, investment in larger or more specialized commercial models—for example Gemini-2.5-flash—is more likely to be justified. A further managerial consideration concerns Explicitness: models that score high on the Explicitness metric list skills verbatim and repeatedly rather than weaving them into readable prose. Whether this affects search-engine visibility or candidate response was not tested in this study so the practical suggestion is confined to reviewing generated postings for repetitive listing of skills before publication. Taken together, the practical recommendation is a test-and-select workflow in which candidate models are validated per domain against the quality dimension that matters most for the intended use case.

5.3. Ethical Implications

Because the generated content is intended for hiring contexts, bias and fairness are material concerns (Bender et al., 2021). Large language models can reproduce and amplify social biases present in their training data, which may surface in AI-generated job postings as gendered or otherwise exclusionary language, uneven representation across demographic groups, or the reinforcement of stereotypes associated with particular occupations. Deploying such content without oversight risks discriminatory outcomes in recruitment (Raghavan et al., 2020). Future work should audit the generated postings for biased or exclusionary language, evaluate fairness across demographic dimensions, and incorporate human review and de-biasing safeguards before any operational use.

6. Conclusions

This research aimed to evaluate synthetically generated job postings produced by large language models based on domain-specific skill areas, comparing them with real-world occupational descriptions. The proposed evaluation methodological framework investigated both linguistic and semantic aspects of generated text by applying multiple complementary metrics, enhancing a fair and systematic comparison among the selected generative models. To minimize the risk of misleading conclusions derived solely from descriptive statistics, a statistical analysis was further conducted to verify and validate the observed experimental results as well as to ensure their quality and reliability. The key findings of the outcomes indicate that no single generative model was superior over the others across all evaluation metrics. Model performance scores were varied depending on the metric under evaluation. In particular, Gemini-2.5-flash and NVIDIA: Nemotron nano achieved the strongest and most comparable overall performance, leading in five and four evaluation-metric cases respectively; NVIDIA: Nemotron nano additionally produced the most plausible output among the examples inspected in the illustrative comparison for the Craft-related domain. This outcome emphasizes the metric-dependent nature of model performance, where better results in specific aspects of text generation occur at the cost of domain areas.
From the domain perspective, the Craft-related domain consistently provided stronger results by the generative models, as examined in both mean value and statistical analysis. This observed trend may result from the lower linguistic complexity and more standardized vocabulary related to craft and trade occupations enhancing more effective content generation. In contrast, domains such as Finance and Healthcare demand specialized and technical language, which may lead generative models to face greater challenges in generating content and consequently result in poor performance in these sectors.
Despite this study providing a comprehensive evaluation framework, it is also subject to various limitations. Initially, the analysis was restricted to a specific set of domains and evaluation metrics, which may not recognize all aspects of text quality or domain expertise. There is also a limitation associated with the statistical analysis. The homogeneity-of-variance assumption underlying the two-way ANOVA was violated for every metric, and no heteroscedasticity-robust two-factor procedure was applied; the F-tests are therefore reported as descriptive rather than exact, supported by a rank-based test on the model factor and by the effect sizes in Table 3. For the same reason no confidence intervals are reported for the effect sizes, since their nominal coverage would not be reliable under the observed heteroscedasticity. A robust two-factor re-analysis, together with a log-transformed analysis of Perplexity and a structured human evaluation of posting quality with independent raters and an inter-rater agreement statistic, are the three most valuable extensions of this work. Two limitations concern the strength of the inferences that can be drawn from the design. The first is prompt sensitivity. All results rest on a single prompt formulation per domain, held constant across the ten models. Holding the prompt fixed is what isolates the model as the source of variation and makes the comparison between models interpretable, but it also means that the rankings reported here are rankings under that prompt. Because the phrasing of an instruction can materially affect the output of a generative model, it cannot be assumed that the ordering of models would be preserved under a differently worded instruction, and in particular a model that responds well to the persona-style framing adopted here may not retain its advantage under a tenser or more schematic instruction. A prompt-sensitivity analysis, in which several paraphrases of each domain prompt are run, and the stability of the ranking is measured, would establish how robust these conclusions are and is the single most informative extension of this work. The second limitation concerns the reporting of variability. The results in Table 1 are cell means, without standard deviations or confidence intervals, so the dispersion of each model within a domain cannot be read from the table and the stability of individual model performance cannot be assessed directly. The statistical analysis in Table 3 partly compensates, since the effect sizes reported there express the proportion of variance attributable to the model, domain and their interaction, and therefore speak to the practical as well as the statistical weight of the differences; but they are not a substitute for per-cell dispersion measures, and future work reporting this benchmark should publish standard deviations and confidence intervals alongside the means. A further limitation concerns the generation configuration: the maximum output-token limit differed substantially across models, from 2048 tokens for Llama-3.1 to more than 100,000 for several others (Table S2.2). Differences in the length and completeness of the generated postings, including those observed in the qualitative comparison, may therefore reflect these configuration ceilings as well as model capability. In particular, the Explicitness metric captures the intensity of verbatim skill enumeration rather than posting quality as such, and none of the five automatic metrics has been validated against human judgements of posting quality; the absence of such a validation, and of a structured human evaluation with independent raters, is a limitation of the present design. In addition, the three domains are not defined at the same level of granularity: “Craft and related trades workers” is an ISCO major group, whereas Finance and Healthcare are broad sectors. This asymmetry partly confounds the finding that the Craft domain is “easier”; besides lower linguistic complexity and more standardized vocabulary, differences in posting volume, breadth, and real-data quality across the three domains could also explain the effect. In addition, the selection of generative models was limited to those that offer free licensing, as including paid models would significantly increase the cost of the research. To be precise about what “free licensing” means operationally, every model in this study was accessed remotely through a hosted cloud API under a free access tier, and none was run through a consumer web interface, through a paid subscription plan, or on local hardware. All generation runs were executed in December 2025. Model identifiers, providers, access routes and execution dates are reported in Table S2.1, and the decoding parameters actually applied to each model—temperature, top-p, maximum output tokens and random seed—in Table S2.2. This uniformity of access route matters for interpretation, since the access level can affect rate limits, context and output ceilings, and in some cases the served variant of a model; because all ten models were accessed on the same terms, differences in output cannot be attributed to differences in access conditions, although the output-token ceilings themselves differ between providers and are discussed as a limitation in Section 6. The selected open-weight and free-tier models were nonetheless chosen to span diverse architectures and parameter scales, so that the comparison still captures the cost–quality trade-offs most relevant to business adoption. Relatedly, each model was queried with the default decoding configuration of its own provider rather than with a single common setting, and a fixed random seed was used throughout. Because provider defaults differ, the temperature and top-p values actually applied vary across models and are reported in Table S2.2. This preserves the out-of-the-box behavior that a practitioner would encounter, but it also means that the models were not run under numerically identical sampling parameters, which limits the strictness of the comparison. Moreover, the utilization of few-shot techniques incorporated in the design of prompts was constrained by computational and financial considerations, while fine-tuning the selected models would require additional resources beyond the scope of the current research. Finally, the construction of ground truth datasets consisted solely of occupation job postings and did not include content related to human resources, which may have limited the applicability of the findings. However, despite the limitations of the current research, the results are very encouraging and indicate high potential for large language models to generate contemporary job postings and, overall, high potential for labor market analysis.
From a business standpoint, this benchmark is best read as a decision-support framework rather than a simple ranking of winners. Current large language models are not yet able to replace human recruiters, but they can act as capable co-pilots in the drafting of job postings. We therefore recommend that organizations adopt a test-and-select strategy, evaluating candidate models on the quality dimension that matters most for the intended use case and validating them separately for each domain before deployment. Whether such adoption shortens time-to-hire, improves applicant quality or yields a positive return on investment was not measured in this study and remains an open question for field research.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/businesses6030043/s1.

Author Contributions

Conceptualization, K.G. and L.A.; methodology, K.G. and L.A.; software, A.A.; validation, K.G. and L.A.; formal analysis, A.A.; investigation, A.A.; data curation, A.A.; writing—original draft preparation, A.A.; writing—review and editing, A.A. and K.G.; visualization, A.A.; supervision, L.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the European Union’s Horizon Europe Framework Programme, SKILLAB project, grant number 101132663.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Appendix A.1. Indicative Generated Jobs for Healthcare and Financial Domains

The descriptions illustrated in Table A1 correspond to job postings related to the Healthcare domain. It is observed that the real job description is more detailed and longer compared to the synthetically generated descriptions by NVIDIA: Nemotron nano and Llama-3.1 respectively. The content generated by models is more concise and role-based, prioritizing direct specifications about the position rather than extended contextual details regarding the organization. The shorter length of generated descriptions may be partially correlated to the small size of the selected models, a choice that was taken under free licensing constraints. For smaller models, it is challenging to maintain coherence over longer texts while keeping the contextual relevance. Moreover, a repetition pattern was observed within the outputs of the models, revealing a possible limitation of the models with variability in context. This pattern may be related to default model settings, specifically with the default temperature parameter, which is responsible for the randomness and creativity of generated output. All models were executed under their respective provider defaults, as described in Section 3 and reported in Table S2.2; the sampling parameters therefore differ between models, and part of the variability observed here may follow from those differences rather than from model capability alone. Despite the strong performance that models have shown in evaluation metrics, the qualitative comparison reveals that generated descriptions are less comprehensive than real-world postings. More specifically, the absence of further and detailed information regarding the organization, work environment, and overall role expectations reveal that even though smaller models demonstrate promising linguistic and semantic alignment with real job descriptions, they are not in a position to fully replace a professionally human-written occupational posting regarding the Healthcare domain.
Table A1. Healthcare domain.
Table A1. Healthcare domain.
Healthcare Domain
Ground TruthBoost your career with Vitalis Medical Metz! Vitalis Medical Metz, temporary—vacation—CDI—CDD recruitment agency, specialized in paramedical, medical and social, is recruiting for its client, an EHPAD, a DE nurse (m/f) in Joeuf. Come and benefit from numerous advantages: 10% paid vacation compensation, 10% end-of-mission compensation, sponsorship bonus (depending on conditions), company mutual insurance, FASTT and CE. As part of your missions as a nurse in the health sector, you will be responsible for: providing hygiene and comfort care to residents, regularly monitoring the state of health of residents, actively contributing to the continuous improvement in the quality of care and the safety of residents, and performing medical procedures in accordance with prescriptions. Prerequisite: you are interested in the wellbeing of residents. Ready to take on exciting challenges? Our Vitalis Medical Metz agency is waiting for you! Profile sought: nursing diploma, ADELI number, registration with the order of nurses. Additional information: type of contract: temporary; working time: full time; salary: negotiable according to profile.
NVIDIA: Nemotron nanoScreening of patients’ dental health, providing cleaning services and preventive care, and educating patients on maintaining oral health.
Llama-3.1Develop chiropractic treatment plans, diagnose problems in the musculoskeletal system, develop a long-term treatment course for disorders in the glandular system.
A similar pattern is also observed in Finance domain descriptions, as displayed in Table A2. Once again, the lack of detailed and longer descriptions in generated postings is evident compared to real-world ones. The same characteristics and constraints of small models also appeared for this domain, as the comparison was made under identical prompt designs and default model settings. The models tend to follow the same direction of direct specification regarding the job role and fully miss the information related to the company environment, corporate culture, and generally what a potential candidate should expect to see on a job posting. Although the quantitative evaluation metrics may provide competitive results in specific areas, the qualitative comparison shows that synthetically generated outputs lack the professional depth and richness in content that is present in real job postings in the Finance domain.
Table A2. Financial domain.
Table A2. Financial domain.
Financial Domain
Ground TruthBasic competencies: payroll management; issuance asset management; asset register monitoring; management of electronic books; classification of expenses; agreement of income–expenses with accounting books; compilation and submission of monthly and other periodic obligations of the company such as VAT, withheld taxes, insurance funds, etc.; participation in the preparation of financial statements and disclosures. Necessary knowledge and skills: degree in economic studies/HEI–TEI; at least two (2) years of experience in a similar position; experience in C class bookkeeping and knowledge of tax legislation; good knowledge and use of ERP software and financial management programs; very good knowledge of MS office; very good knowledge of English. The company offers employment opportunities in an excellent working environment, with a competitive salary package.
NVIDIA: Nemotron nanoAnalyzes financial data to provide insights into company performance, budgeting, and forecasting. Supports decision-making through detailed financial reports and trend analysis.
Llama-3.1Develop and implement training programs to enhance employee skills and knowledge, aligning with business objectives and regulatory requirements.

Appendix A.2. Construction of the Domain Skill Sets

All skills used as generation inputs were drawn from a single taxonomy release, ESCO v1.2.0, so that skill labels are internally consistent across the three domains. Both principal branches of the taxonomy were used: the skills and competences pillar and the knowledge pillar. Including the knowledge pillar is material for this study, because a substantial part of what distinguishes Finance and Healthcare postings from Craft postings is domain knowledge—regulatory frameworks, accounting standards, clinical terminology—rather than task-level skills, and excluding it would have systematically impoverished precisely the two domains that proved hardest to generate.
The construction procedure was identical for each domain. First, the occupations belonging to the domain were identified in ESCO. Second, all skills and knowledge concepts associated with those occupations were retrieved. No distinction was made between essential and optional associations: every concept linked to an occupation of the domain was retained, on the grounds that a job posting may legitimately reference either. Third, the retrieved URIs were mapped to human-readable labels using the ESCO reference dictionaries described in Section 3, and the resulting labels were assembled into the domain skill set that was supplied to the models in the prompt.
No deduplication rule and no cross-domain overlap-removal procedure were applied to the assembled skill sets. Separation between the three domains therefore rests entirely on the a priori selection of substantively distinct occupational areas, and not on any post hoc filtering step. Two consequences should be noted explicitly. A skill or knowledge concept that ESCO associates with occupations in more than one of the three domains will appear in more than one skill set, so the sets are not guaranteed to be disjointed; and because duplicate labels were not collapsed within a set, a concept reachable through several occupations of the same domain may be represented more than once. Neither effect biases the comparison between models, since every model received exactly the same skill set for a given domain, but both weaken the claim that the three domains constitute fully independent skill environments, and the latter may inflate the Explicitness metric for domains whose skill sets contain more repeated labels. This is acknowledged as a limitation in Section 6.
Table A3. Composition of the three domain skill sets and of the corresponding reference datasets.
Table A3. Composition of the three domain skill sets and of the corresponding reference datasets.
DomainTaxonomic Scope (ESCO v1.2.0)ESCO OccupationsSkills and Knowledge ConceptsReference Job Postings
FinanceISCO-08 major group 1Finance managersAll skills extracted from the retrieved job postings15,000
HealthcareISCO-08 major group 3Health associate professionalsAll skills extracted from the retrieved job postings15,000
Craft and related trades workersISCO-08 major group 70 (major group)All skills extracted from the retrieved job postings20,000
Total 50,000
The provenance of these postings is as follows. All records were retrieved from the SKILLAB API, which aggregates online vacancy data for the European labor market, and were processed with the preprocessing pipeline described in Section 3: removal of records with missing skill fields, mapping of ESCO URIs to human-readable labels, discarding of non-essential attributes, lowercasing and removal of HTML markup, and finally filtering by domain. The data were retrieved from the SKILLAB API and spanned the EU 27 countries, for the entirety of 2025, for the three domains.

References

  1. Abdelhay, S., AlTalay, M. S. R., Selim, N., Altamimi, A. A., Hassan, D., Elbannany, M., & Marie, A. (2025). The impact of generative AI (ChatGPT) on recruitment efficiency and candidate quality: The mediating role of process automation level and the moderating role of organizational size. Frontiers in Human Dynamics, 6, 1487671. [Google Scholar] [CrossRef] [Scilit]
  2. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021, March 3–10). On the dangers of stochastic parrots: Can language models be too big? [Conference paper]. 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ‘21) (pp. 610–623), Toronto, ON, Canada. [Google Scholar]
  3. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, A., Herbert-Voss, A., Krueger, H., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., … Amodei, D. (2020, December 6–12). Language models are few-shot learners [Conference paper]. Advances in Neural Information Processing Systems 33 (NeurIPS 2020) (pp. 1877–1901), Virtual. [Google Scholar]
  4. Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942. [Google Scholar] [CrossRef] [Scilit]
  5. Cao, Y., Li, S., Liu, Y., Yan, Z., Dai, Y., Yu, P. S., & Sun, L. (2023). A comprehensive survey of AI-generated content (AIGC): A history of generative AI from GAN to ChatGPT. arXiv. [Google Scholar] [CrossRef] [Scilit]
  6. Cappelli, P. H. (2015). Skill gaps, skill shortages, and skill mismatches: Evidence and arguments for the United States. ILR Review, 68(2), 251–290. [Google Scholar]
  7. Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science, 384(6702), 1306–1308. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Forsythe, E., Kahn, L. B., Lange, F., & Wiczer, D. (2020). Labor demand in the time of COVID-19: Evidence from vacancy postings and UI claims. Journal of Public Economics, 189, 104238. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Jelinek, F., Mercer, R. L., Bahl, L. R., & Baker, J. K. (1977). Perplexity—A measure of the difficulty of speech recognition tasks. The Journal of the Acoustical Society of America, 62(S1), S63. [Google Scholar] [CrossRef] [Scilit]
  10. Khaouja, I., Kassou, I., & Ghogho, M. (2021). A survey on skill identification from online job ads. IEEE Access, 9, 118134–118153. [Google Scholar] [CrossRef] [Scilit]
  11. Kruskal, W. H., & Wallis, W. A. (1952). Use of ranks in one-criterion variance analysis. Journal of the American Statistical Association, 47(260), 583–621. [Google Scholar] [CrossRef] [Scilit]
  12. Kusner, M. J., Sun, Y., Kolkin, N. I., & Weinberger, K. Q. (2015, July 6–11). From word embeddings to document distances [Conference paper]. 32nd International Conference on Machine Learning (ICML 2015) (pp. 957–966), Lille, France. [Google Scholar]
  13. le Vrang, M., Papantoniou, A., Pauwels, E., Fannes, P., Vandensteen, D., & De Smedt, J. (2014). ESCO: Boosting job matching in Europe with semantic interoperability. Computer, 47(10), 57–64. [Google Scholar] [CrossRef] [Scilit]
  14. Magron, A., Dai, A., Zhang, M., Montariol, S., & Bosselut, A. (2024, March 17–22). JobSkape: A framework for generating synthetic job postings to enhance skill matching [Conference paper]. First Workshop on Natural Language Processing for Human Resources (NLP4HR 2024) (pp. 43–58), St. Julian’s, Malta. [Google Scholar]
  15. Mahjoub, A., & Kruyen, P. M. (2021). Efficient recruitment with effective job advertisement: An exploratory literature review and research agenda. International Journal of Organization Theory & Behavior, 24(2), 107–125. [Google Scholar] [CrossRef] [Scilit]
  16. Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020, January 27–30). Mitigating bias in algorithmic hiring: Evaluating claims and practices [Conference paper]. 2020 Conference on Fairness, Accountability, and Transparency (FAT* ‘20) (pp. 469–481), Barcelona, Spain. [Google Scholar]
  17. Reimers, N., & Gurevych, I. (2019, November 3–7). Sentence-BERT: Sentence embeddings using siamese BERT-networks [Conference paper]. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992), Hong Kong, China. [Google Scholar]
  18. Tausczik, Y. R., & Pennebaker, J. W. (2010). The psychological meaning of words: LIWC and computerized text analysis methods. Journal of Language and Social Psychology, 29(1), 24–54. [Google Scholar] [CrossRef] [Scilit]
  19. Tuttle, L., & Critchlow, K. (2025). Digital transformation in talent acquisition: Modern approaches to recruitment and selection. International Journal of Research in Human Resource Management, 7(1), 351–357. [Google Scholar] [CrossRef] [Scilit]
  20. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017, December 4–9). Attention is all you need [Conference paper]. Advances in Neural Information Processing Systems 30 (NeurIPS 2017) (pp. 5998–6008), Long Beach, CA, USA. [Google Scholar]
  21. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., … Wen, J. R. (2023). A survey of large language models. arXiv. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Methodology stages.
Figure 1. Methodology stages.
Businesses 06 00043 g001
Figure 2. Heatmap results of best performing models per domain.
Figure 2. Heatmap results of best performing models per domain.
Businesses 06 00043 g002
Figure 3. Best-performing model trade-offs.
Figure 3. Best-performing model trade-offs.
Businesses 06 00043 g003
Figure 4. Domain facilitation effect across all models.
Figure 4. Domain facilitation effect across all models.
Businesses 06 00043 g004
Table 1. Results of evaluation metrics used to assess model performance (C.S. denotes Cosine Similarity; arrows indicate the optimization direction for the four directional metrics—↑ higher is better, ↓ lower is better. No arrow is attached to Explicitness: as explained in Section 3, that metric is not monotonically directional, and its values are reported descriptively rather than as a quality ranking).
Table 1. Results of evaluation metrics used to assess model performance (C.S. denotes Cosine Similarity; arrows indicate the optimization direction for the four directional metrics—↑ higher is better, ↓ lower is better. No arrow is attached to Explicitness: as explained in Section 3, that metric is not monotonically directional, and its values are reported descriptively rather than as a quality ranking).
ModelSkill DomainPerplexity (↓)LIWC-Style Divergence (↓)Mean C.S. (↑)WMD (↓)Explicitness
Deepseek-r1Finance132.680.400.4871.217.34
Deepseek-r1Health91.400.490.2141.247.30
Deepseek-r1Craft38.830.420.4201.1311.57
Gemini-2.5-flashFinance47.710.420.5201.1629.03
Gemini-2.5-flashHealth51.750.290.5651.1514.69
Gemini-2.5-flashCraft53.510.300.4921.1612.17
Gemma3Finance148.070.400.5131.2311.47
Gemma3Health97.740.410.3741.239.20
Gemma3Craft115.140.450.5131.2010.10
GPT-OSSFinance213.300.380.4271.2210.93
GPT-OSSHealth134.490.410.3191.228.19
GPT-OSSCraft151.510.470.4481.217.92
Grok-4.1Finance398.870.320.2971.1914.97
Grok-4.1Health462.820.350.5221.209.18
Grok-4.1Craft723.900.390.4581.1913.34
Hermes-3Finance60.510.410.4611.2010.31
Hermes-3Health46.560.410.3231.189.19
Hermes-3Craft77.100.410.5901.188.73
Llama-3.1Finance95.900.400.3361.2011.52
Llama-3.1Health171.460.480.2591.129.29
Llama-3.1Craft2680.750.630.4621.195.75
MistralFinance136.810.260.3241.2021.73
MistralHealth113.370.380.3391.2110.42
MistralCraft51.050.360.3811.1212.58
NVIDIA: Nemotron nanoFinance54.150.470.5541.207.00
NVIDIA: Nemotron nanoHealth116.670.400.3561.228.68
NVIDIA: Nemotron nanoCraft34.070.430.3931.1029.49
Qwen-2.5Finance52.500.350.4081.1711.76
Qwen-2.5Health40.860.520.3651.1811.64
Qwen-2.5Craft463.810.480.4331.188.55
Table 2. Best-performing model per evaluation metric and domain, derived directly from Table 1. The final column reports how many of the five metrics each domain’s leading model wins. Because the Explicitness metric is not monotonically directional (Section 3), no optimization arrow is attached to it and the model listed in that column is not identified as the best performer: the entry records only which model produced the lowest verbatim-enumeration intensity, and is reported for description. Accordingly, the metric-win counts in the final column are computed over the four directional metrics only, and Explicitness is excluded from them. Metric-win counts are descriptive only and are not accompanied by a measure of uncertainty.
Table 2. Best-performing model per evaluation metric and domain, derived directly from Table 1. The final column reports how many of the five metrics each domain’s leading model wins. Because the Explicitness metric is not monotonically directional (Section 3), no optimization arrow is attached to it and the model listed in that column is not identified as the best performer: the entry records only which model produced the lowest verbatim-enumeration intensity, and is reported for description. Accordingly, the metric-win counts in the final column are computed over the four directional metrics only, and Explicitness is excluded from them. Metric-win counts are descriptive only and are not accompanied by a measure of uncertainty.
DomainPerplexity (↓)LIWC-Style Divergence (↓)Mean C.S. (↑)WMD (↓)Explicitness (Lowest Value, Descriptive)Most Top Scores
CraftNVIDIA: Nemotron nanoGemini-2.5-flashHermes-3NVIDIA: Nemotron nanoLlama-3.1NVIDIA: Nemotron nano (2)
FinanceGemini-2.5-flashMistralNVIDIA: Nemotron nanoGemini-2.5-flashNVIDIA: Nemotron nanoGemini-2.5-flash (2)
HealthcareQwen-2.5Gemini-2.5-flashGemini-2.5-flashLlama-3.1Deepseek-r1Gemini-2.5-flash (2)
Table 3. Results of statistical tests (classical two-way ANOVA with the Kruskal–Wallis robustness check on the model factor). Each model × domain cell contains n = 100 generated postings, giving N = 3000. In the ANOVA rows the model, domain and interaction effects carry 9, 2 and 18 numerator degrees of freedom respectively, with 2970 residual degrees of freedom; the Kruskal–Wallis statistic on the model factor carries 9 degrees of freedom. The Statistic column reports the test statistic and the p-Value column the corresponding significance level. The Effect Size column reports partial eta-squared for the ANOVA rows and epsilon-squared for the Kruskal–Wallis rows, both computed from the statistics and degrees of freedom shown. Because the homogeneity-of-variance assumption is violated for every metric, these effect sizes and the accompanying p-values are descriptive; no confidence intervals are reported, as their nominal coverage would not be reliable under the observed heteroscedasticity (Section 6).
Table 3. Results of statistical tests (classical two-way ANOVA with the Kruskal–Wallis robustness check on the model factor). Each model × domain cell contains n = 100 generated postings, giving N = 3000. In the ANOVA rows the model, domain and interaction effects carry 9, 2 and 18 numerator degrees of freedom respectively, with 2970 residual degrees of freedom; the Kruskal–Wallis statistic on the model factor carries 9 degrees of freedom. The Statistic column reports the test statistic and the p-Value column the corresponding significance level. The Effect Size column reports partial eta-squared for the ANOVA rows and epsilon-squared for the Kruskal–Wallis rows, both computed from the statistics and degrees of freedom shown. Because the homogeneity-of-variance assumption is violated for every metric, these effect sizes and the accompanying p-values are descriptive; no confidence intervals are reported, as their nominal coverage would not be reliable under the observed heteroscedasticity (Section 6).
MetricTestEffectStatisticp-ValueSignificantEffect Size
PerplexityTwo-way ANOVAModelF(9, 2970) = 0.990.44No0.003
PerplexityTwo-way ANOVADomainF(2, 2970) = 2.240.11No0.002
PerplexityTwo-way ANOVAInteractionF(18, 2970) = 1.300.18No0.008
PerplexityKruskal–WallisModelH(9) = 1670.39<0.001Yes0.556
LIWC-style divergenceTwo-way ANOVAModelF(9, 2970) = 100.77<0.001Yes0.234
LIWC-style divergenceTwo-way ANOVADomainF(2, 2970) = 684.30<0.001Yes0.315
LIWC-style divergenceTwo-way ANOVAInteractionF(18, 2970) = 56.20<0.001Yes0.254
LIWC-style divergenceKruskal–WallisModelH(9) = 505.07<0.001Yes0.166
Mean C.S.Two-way ANOVAModelF(9, 2970) = 224.52<0.001Yes0.405
Mean C.S.Two-way ANOVADomainF(2, 2970) = 240.96<0.001Yes0.140
Mean C.S.Two-way ANOVAInteractionF(18, 2970) = 90.54<0.001Yes0.354
Mean C.S.Kruskal–WallisModelH(9) = 1087.62<0.001Yes0.361
WMDTwo-way ANOVAModelF(9, 2970) = 121.52<0.001Yes0.269
WMDTwo-way ANOVADomainF(2, 2970) = 106.00<0.001Yes0.067
WMDTwo-way ANOVAInteractionF(18, 2970) = 61.90<0.001Yes0.273
WMDKruskal–WallisModelH(9) = 716.06<0.001Yes0.236
ExplicitnessTwo-way ANOVAModelF(9, 2970) = 207.21<0.001Yes0.386
ExplicitnessTwo-way ANOVADomainF(2, 2970) = 267.14<0.001Yes0.152
ExplicitnessTwo-way ANOVAInteractionF(18, 2970) = 138.13<0.001Yes0.456
ExplicitnessKruskal–WallisModelH(9) = 874.89<0.001Yes0.290
Table 4. Indicative job descriptions of Craft-related domain.
Table 4. Indicative job descriptions of Craft-related domain.
Craft-Related Domain
Ground TruthThe main responsibilities include diagnosis and restoration of electrical equipment; supervision and restoration of building and plumbing installations; operation and maintenance of electrical, heating, cooling, machinery and forklifts; ability to work in shifts; professional qualifications and experience; an electrician’s or assistant electrician’s license; a TEI or HEI degree in a related subject is not required but will be considered; a driver’s license; good of PC and MS office experience from 3 to 5 years in a similar position. Skills critical: thinking ability to make decisions, results oriented, creativity, teamwork and co-operation. Satisfactory remuneration offered; dynamic work environment with excellent training and development prospects.
NVIDIA: Nemotron nanoJoin SolarWise Energy Solutions, a rapidly growing company specializing in renewable energy systems. You will install and maintain photovoltaic panels on residential and commercial properties, ensuring optimal performance and compliance with safety standards. Responsibilities include site assessments, panel installation, wiring, commissioning, and routine maintenance. You must have a technical background in electrical systems, experience with solar technology, and a strong commitment to sustainability. B license required.
Llama-3.1Lead the design and development of complex electrical systems for commercial and industrial applications. This role involves conceptualization, detailed circuit design, component selection, and overseeing the prototyping and testing phases. You will work closely with cross-functional teams to ensure designs meet performance, cost, and reliability targets, while adhering to relevant industry standards and regulations. A strong understanding of power electronics and control systems is essential.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Adam, A.; Georgiou, K.; Angelis, L. Benchmarking Generative AI Models for Skill-Aligned Job Posting Generation: A Multi-Domain Comparative Evaluation. Businesses 2026, 6, 43. https://doi.org/10.3390/businesses6030043

AMA Style

Adam A, Georgiou K, Angelis L. Benchmarking Generative AI Models for Skill-Aligned Job Posting Generation: A Multi-Domain Comparative Evaluation. Businesses. 2026; 6(3):43. https://doi.org/10.3390/businesses6030043

Chicago/Turabian Style

Adam, Alexandros, Konstantinos Georgiou, and Lefteris Angelis. 2026. "Benchmarking Generative AI Models for Skill-Aligned Job Posting Generation: A Multi-Domain Comparative Evaluation" Businesses 6, no. 3: 43. https://doi.org/10.3390/businesses6030043

APA Style

Adam, A., Georgiou, K., & Angelis, L. (2026). Benchmarking Generative AI Models for Skill-Aligned Job Posting Generation: A Multi-Domain Comparative Evaluation. Businesses, 6(3), 43. https://doi.org/10.3390/businesses6030043

Article Metrics

Back to TopTop