The method applied in this investigation takes its inspiration from a recent article [
6] (hereinafter referred to as Birhane et al.), which applies a discourse analysis approach to extract statements of researcher ‘value commitments’ within a corpus of papers from the field of machine learning. Discourse analysis distinguishes itself from other applications of corpus linguistics in its goal to “shed light on how speakers indicate their semantic intentions, how hearers interpret what they hear, and the cognitive abilities that underlie human symbolic use” [
7]. The corpus underlying this analysis is not large, focusing on 100 highly cited papers from two prestigious conference proceedings (NeurIPS and ICML), and its method was labour intensive, with humans annotating and manually encoding the abstract, introduction, discussion, and conclusion sections of each paper in the search for statements of researcher value positionings and ‘justificatory chains,’ that is the lines of reasoning by which each paper makes its case as a contribution to scholarship. The work ultimately concludes that the underlying values expressed within this corpus are out of step with social goods, following instead a tacit political and social agenda that centralises technological power and maintains existing hierarchies.
What is interesting about this work is the manner in which it is able to extract tacit positioning from discourse that was not necessarily intended to be read for these signals. It was this desire not to be led by assumptions that guided Birhane et al. to adopt the labour-intensive dual human encoding method they used to indicate justificatory claims, a slightly different goal from the investigation that follows, which is more concerned with how research is perceived. And yet, taking Birhane et al.’s questions and dataset as a starting point allows for a highly instructive set of comparisons to be made, in particular, as the dataset underpinning their work has been made openly available for further analysis. What follows in this paper, therefore, presents some of the methods applied by Birhane et al. and compares them with equivalent corpora drawn from both a traditional humanities (taking literary studies as an example) and a digital humanities journal. The resulting evidence establishes a framework for clarifying the nuances of these different epistemic cultures, and the nature of their current and potential future state of convergence, from the perspective of discourses related to values, perceptions and approaches, highlighting, in particular, the divergences among these fields. This paper transforms the original authors’ single-discipline dataset into a cross-disciplinary one, adapting the original method in the following ways:
2.1. Moving from a Single to a Comparative Corpus
In Birhane et al.’s original paper, the curation of the data was based upon the notion of papers being influential because they are highly cited, and being presented at certain high-profile conferences, which are indicative of community influence in the computer science disciplines. The use of citations as a proxy for influence does not operate in the same way in the arts and humanities, however, and has been widely criticised when deployed as such (see, for example, [
8,
9,
10]). This is not to say that there are not more and less prominent, established journals, however, and ones which, in their own way, indicate a similar status of representing the voice of a research community. Conferences do not have this status across the disciplines, however, but what humanities disciplines have instead are conferences and associated journals that are published by the large membership associations. The editorial processes of these organs, which are responsive to the voices of the members, can on this basis be understood as a good proxy for community norms and practices.
On this basis, my comparator corpora have been drawn from two long-running and respected journals, which may be said to speak for the norms of the community due to their positioning as the voices of prominent subject associations. The Publications of the Modern Language Association (established in 1884) is a community organ shared by the over 20,000 members of the Modern Language Association, an organisation that presents itself as a leading advocate for the humanities in general, but more specifically for the study of languages, literatures, and culture. Just as ML research stands in as a proxy in Birhane et al.’s work for the wider AI community, the MLA can stand as a coherent representative of a specific sub-community with broad representation in the wider field of the humanities.
While the challenge with the humanities lies in the presence of a number of established and distinct subfields, finding the ‘voice’ of the digital humanities presents a very different, if no less complex, problem. Kathleen Fitzpatrick’s early and inclusive definition of DH as “the humanities, done digitally” (an extension of the preexisting theory/practice divide long known in the humanities) [
11]. But the simplicity of this formulation overlays a significant body of work that proposes myriad different, sometimes contradictory, definitions of the field, as seen for example in [
12] and in the very existence of a long-running and popular series of volumes known as the Debates in the Digital Humanities (edited by Matthew Gold and Lauren Klein). In spite of this heterogeneity, for the digital humanities, there is an equivalent journal to the PMLA overseen by a prominent subject association in Digital Scholarship in the Humanities (DSH). First published in 1986 (as Literary and Linguistic Computing), the journal has long served as the outlet of the Alliance of Digital Humanities Organisations (ADHO), an umbrella organisation representing over a dozen regional DH associations. As both of these journals are registered in the Web of Science database, abstracts for a similarly sized sample of recently published (2022–2023) articles could be assembled into parallel corpora to be used alongside that of Birhane et al. An overview of the size and complexity of the three corpora appears in
Table 1 below.
The papers were selected as a continuous run, starting in 2022 and ending when the desired number of abstracts (+/− 3) had been reached. The of slight variation in the number of entries in each corpus represents the best possible equivalence between the three sources, balancing overall corpus size and number of entries with the desire to include the full contents of any release batches included in the sample. The data was inspected to remove possible biasing factors, such as special issues or reviews, and the full metadata and abstract text extracted into an Excel spreadsheet, where duplicates were removed, missing data (such as publication dates) added, and any obvious errors in spelling or data organisation (such as data in the wrong columns) corrected. From here, the abstract text was copied into a .txt file able to be processed by the corpus query software package being used.
These journals cannot be understood as wholly representative of the wider fields they are drawn from, but as illustrative proxies that provide a useful and good, but not perfect, representation of wider trends. Just as Birhane et al. abstract from data reflecting the Machine Learning community to wider conclusions about AI in general, these corpora also draw from snapshots of humanities and digital humanities research activity, with a bias toward the English language and Northern/Western perspectives, toward the specific editorial policies and procedures of the journals, and, in particular, the decision to feature a subject association anchored in literary studies to represent the humanities. This is not to say that these proxies do not lose something with their essentially metonymical nature, taking a part as representative of a much larger whole, and perhaps missing wider trends in the broader fields as we tend to conceptualise them. It should be noted, however, that the first real challenge here lies in the ease with which we use terms like ‘artificial intelligence’ and ‘the humanities’ in a way that seems to describe something cohesive and clear, but which actually may incorporate a very wide range of practices. Future research might add further corpora, for example drawn from the proceedings of the American Historical Association, or the journals of some of the regional subject associations that operate under the umbrella of the European Association for the Digital Humanities. This approach, however, would potentially introduce further biases and complexities, while also not abiding by the desire to mirror the original structure presented by Birhane et al.. For this reason, in spite of their limitations, these two journals were ultimately chosen as the best possible matches for the goals of this study. The second hard challenge related to the degree to which these outlets can be seen as comparable, seeing as two of them publish papers submitted directly, whereas the original dataset is comprised of conference papers. My response to this potential critique is twofold: first, in fact, a conference paper in the humanities would be likely to be far more different from a computer science conference paper, as humanities conferences tend to accept or reject not upon full papers, but upon short abstracts. The greater comparability in process and form is reflected in the choices made for this analysis, in particular as pertains to critical aspects such as selection criteria and depth of peer review. Second, we must always remember that investigations of interdisciplinary work require that we suspend some more exacting conceptualisations of comparability: disciplines do, to some extent, inhabit their own epistemic cultures, and to compare between them always involves some flattening of specificities. If we allow this perception of differences to dominate, however, then we close doors to methods that might enhance our understanding of how knowledge creation ecosystems might continue to evolve in conversation with each other.
First, please clarify the comparability of the three corpora, particularly differences in publication type, venue, selection criteria, and citation status.
2.3. Revisiting Goals and Reviving a Quantitative Approach
Finally, I choose to deploy a mixed-methods, but largely automated analysis workflow, rather than the qualitative and human-encoded one employed by Birhane et al. Such an approach had been considered and rejected by Birhane et al.’s team out of a desire to ensure that the values that were identified were not led by confirmation bias. A number of factors made such an approach a more reasonable choice for this comparative study, however. First of all, a focus on habits of discourse, rather than more complex statements of values, meant that specific linguistic features could be more telling than they might have been for Birhane et al. Furthermore, observation of the annotated data produced by Birhane et al. showed that the value positions extracted could generally be mapped onto particular discrete linguistic features, supporting an analysis that is evidence-based, but ultimately more descriptive than inferential. Reading through the justificatory chains identified in the original dataset demonstrated some very clear patterns. For example, in the examples featured in their article, the value of “efficiency” was always flagged with specific use of this word. Other values, such as novelty, hewed to relatively narrow ranges of expression. Given the focus on discourse rather than underlying values, the automated approach rejected by Birhane et al. could play a larger role here, particularly due to the presence of specific intention statements made by the authors, clearly indicated with the use of the personal pronoun “we” (as in “in this paper, we show …”).
The basis of my analysis then became precisely this: to extract computationally from the three parallel abstract corpora such specific statements of authorial intent, and to compare them in their form, but also in terms of the implications of the words used to characterise the epistemic claims being made. The workflow implemented for this was relatively straightforward, using Microsoft Excel to clean and verify data, and the LancsBox [
19] toolkit to apply tags (grammatical and semantic), process and analyse data samples largely using KWIC queries to find and identify the function of single word agency statements and multi-word characterisations of the actions of these agents (active verb following the noun or pronoun reflecting the perceived agent). This feature selection was verified manually, with any ambiguous cases (such as the humanistic use of the word ‘we,’ discussed below) captured and considered for their weight within the interpretive framework. As the corpus was small, all manual encoding was carried out by the author.
What follows below are the observations facilitated by this analysis process and its results, grouped under four different headings: first, a consideration of the values found by Birhane et al. and their lack of resonance in the comparator corpora; second, the nature of the agent making an epistemic claim in a given paper from a given field; third, the discourses of the justificatory chain, in particular, the nature of the action of research being portrayed and finally, the specific positionality of the digital humanities in terms of its blending of disciplinary discourses and cultures.