2.2. Querybank198 Pair Construction and Development-to-Held-Out Audio-Disjoint Evaluation
Querybank198 is a derived evaluation protocol constructed from publicly obtainable Tibetan speech resources, including TIBMD@MUC, XBMU-AMDO31, Tibetan Greetings, and NICT-Tib1 [
24,
25,
26,
27]. These resources cover Amdo, Kham, and U-Tsang/Lhasa-related speech recorded under different corpus, speaker, and acoustic conditions. Querybank198 releases protocol metadata and indexed pair files rather than redistributing the source audio. The protocol materials can be mapped to locally obtained copies of the four source resources through the provided path-mapping configuration.
The protocol distinguishes the unit used for keyword-spotting evaluation from the unit used for split-independence auditing. The evaluation unit is an indexed query–speech pair. A single speech recording can be evaluated against multiple written queries and can therefore generate multiple positive or negative pairs. Consequently, the number of query–speech pairs is not equivalent to the number of independent speech recordings. The recording unit used for the independence audit is the complete referenced waveform-audio (WAV) file. The approximately 2.0 s acoustic windows used by Tibetan-PASEM are generated from these referenced recordings during feature construction and are not treated as independent stored recordings or split-assignment units.
The original Querybank198 index contained 198 written Tibetan query entries and 214,644 indexed query–speech pairs: 127,536 pairs in seen_train, 28,848 in seen_cv, 43,026 in seen_test, 2706 in unseen_test, and 12,528 in the legacy hardest_test file. The seen_train partition provides model-training pairs for 99 training-exposed query forms, while seen_cv uses the same query role for validation and operating-point selection. The seen_test condition evaluates training-exposed query forms on held-out speech, and unseen_test evaluates 79 written query forms excluded from model training after coverage filtering. After development-to-held-out audio correction, unseen_test contains 2697 indexed pairs, including 451 positive and 2246 negative pairs. This corresponds to an average of 5.71 positive pairs per held-out query form, although the pair distribution is not assumed to be uniform across queries.
The legacy hardest_test file comprises 20 of the 99 training-exposed query forms and evaluates them against a targeted confusable-negative construction. In seen_train, these 20 query forms account for 38,240 query–speech pairs, including 9560 positive pairs, representing 30.0% of all seen_train pairs. This corresponds to an average of 478 positive pairs per query, compared with 322.06 positive pairs per query across all 99 seen queries. The revised protocol therefore terms this evaluation the confusable-negative condition and interprets it as a targeted diagnostic of negative-set composition rather than an ordered difficulty tier. These evaluation roles follow user-defined and open-vocabulary KWS settings in which familiar-query detection, held-out written-query transfer, and query-specific false-trigger behavior may differ [
6,
7,
8,
9,
10,
11,
12,
13,
15].
The protocol is not source-corpus-disjoint. Recordings from the same four public resources may contribute to different query roles so that query-role comparisons are not confounded with a complete change of corpus. Split independence is instead enforced across the development-to-held-out boundary at the recording-content, source-qualified speaker, and derived session-proxy levels. The development partitions comprise seen_train and cleaned seen_cv. The held-out evaluation conditions comprise seen_test, unseen_test, and the confusable-negative condition.
Recording-content independence was audited using canonical file paths and content hashes. Complete-file Secure Hash Algorithm 256 (SHA-256) values were computed for all 91,988 referenced WAV paths. Decoded pulse-code modulation (PCM) SHA-256 values were additionally computed where PCM decoding was supported, allowing files with different names, paths, or container headers but identical decoded acoustic samples to be identified. Files for which decoded-PCM hashing was unavailable remained covered by complete-file hashing and were recorded separately in the audit manifest. Two referenced recordings were assigned to the same acoustic-identity group when either their complete-file hashes or decoded-PCM hashes were identical.
The corrected protocol retains seen_train unchanged. From seen_cv, all query–speech pairs associated with an acoustic-identity group present in seen_train were removed. From seen_test, unseen_test, and the confusable-negative condition, all pairs associated with an acoustic identity present in seen_train or cleaned seen_cv were removed. Every query–speech pair associated with an excluded recording was removed jointly, irrespective of its query identifier or pair label. This procedure removed 1516 pairs and produced a 213,128-pair evaluation: 127,536 pairs in seen_train, 27,577 in seen_cv, 42,843 in seen_test, 2697 in unseen_test, and 12,475 in the confusable-negative condition.
Speaker independence was audited using source-qualified speaker identifiers formed by combining the source-resource identifier with the corpus-specific speaker identifier. No source-qualified speaker is shared between seen_train or cleaned seen_cv and any held-out evaluation condition. The four source resources do not provide a consistent common recording-session identifier. Session behavior was therefore audited using conservative speaker-scoped session proxies derived from explicit corpus path structure, recording-condition folders, dated recording groups, and corpus subfolders. These proxies also show zero overlap across the development-to-held-out boundary. The corrected protocol is recording-content-disjoint under the complete-file and decoded-PCM identity criteria, speaker-disjoint, and proxy-session-disjoint across the development-to-held-out boundary. Official recording-session independence is not inferred beyond the metadata supplied by the original resources.
The three held-out conditions partially reuse held-out recordings and speakers. They are therefore interpreted as correlated condition-specific evaluations of seen-query detection, unseen written-query transfer, and confusable-negative behavior rather than as three statistically independent audio test sets. This held-out-to-held-out reuse does not cross the development-to-held-out boundary. Threshold selection and fixed held-out application follow the policy defined in
Section 2.8 [
10,
12,
15].
Figure 1 summarizes the original pair index, the recording-content correction and parallel speaker/session-proxy audits, and the corrected fixed-threshold evaluation.
The complete removal manifest, whole-file and decoded-PCM hash records, original and cleaned pair indexes, speaker and session-proxy overlap matrices, threshold-selection records, and score-to-pair alignment audits are provided in the
Supplementary Materials. These materials support reconstruction of the Querybank198 protocol and the corrected development-to-held-out evaluation when the four source speech resources are obtained locally [
24,
25,
26,
27,
33]. The released components and their supported reproduction roles are summarized in
Table 2.
Protocol reconstruction consists of local path configuration, pair-index alignment, acoustic-identity correction, and score-level fixed-threshold evaluation. The release supports protocol construction, overlap auditing, threshold selection, and metric recalculation. Complete end-to-end model-training code is outside the current release scope.
2.3. Query Evidence Representation
Each Tibetan written query is represented by three complementary evidence views. The graphemic view g(q) preserves the normalized written Tibetan form. The approximate phonological view p(q) is obtained from inherited lexical or syllable-level entries when available and from a rule-based Wylie-style fallback when no inherited entry is available. The dialect-related feature vector d(q) summarizes conservative syllable- and pronunciation-related cues, including syllable count, onset/rime/coda-like patterns, syllable-structure indicators, and available Amdo, U-Tsang/Lhasa-related, or Kham source indicators. This design follows phoneme- and phonology-aware KWS, where pronunciation-related evidence can help distinguish short and acoustically confusable keywords [
9,
12,
15] while reflecting the dialectal diversity of Tibetic speech varieties [
28,
29].
The query encoder contains a character-token mean encoder, a phonological-token mean encoder, and a linear projection for handcrafted dialect-related features. The three views are concatenated and projected into a shared query evidence embedding:
where
and
encode graphemic and approximate phonological evidence,
projects dialect-related features, and
maps the concatenated evidence into the shared matching space.
The three query views are used as structured query-side matching evidence. The graphemic view preserves the normalized written form, while approximate phonological cues add pronunciation-related information from inherited lexical entries, syllable-level entries, and rule-based Wylie-style fallback mappings. The dialect-related feature vector summarizes conservative syllable and source indicators.
Table 3 specifies the source category used for each approximate phonological representation. This design fits the low-resource setting, where complete dialect-specific pronunciation dictionaries are unavailable, but lightweight pronunciation metadata can improve discrimination among short and confusable written queries [
9,
12,
15,
24,
25,
26,
27,
28,
29].
2.5. Latent-Window Probabilistic Evidence Matching
Tibetan-PASEM uses a latent-window evidence-matching formulation. The latent variable Z denotes the candidate acoustic window that contains the written query. The model estimates local query–window evidence, summarizes the window-level matching logits, and produces the final pair-level posterior through top-ranked local verification. This design matches the Querybank198 supervision format, where pair-level labels are available without manually annotated keyword boundaries.
Using the acoustic-window embedding hi defined in
Section 2.4 and the query embedding eq defined in
Section 2.3, the model estimates a local query–window match probability:
where
denotes elementwise multiplication, [;] denotes concatenation, and
measures the local compatibility between the written query and the
acoustic window. The keyword location remains latent because Querybank198 provides pair-level labels without manually annotated keyword boundaries.
The matcher aggregates the window-level logits through normalized log-sum-exp:
where
is the candidate-window set and
is the pre-sigmoid matching logit. This corresponds to
= 1.0 in the generalized temperature-controlled formulation and provides pair-level matcher supervision without keyword-boundary annotations.
For local verification, the candidate windows are ranked by their matching logits, and the five highest-scoring windows are retained. Each selected window is represented by its acoustic embedding and corresponding matching logit. These top-five representations are concatenated with the query-side feature vector d(q):
The local verifier is a multilayer perceptron with two hidden layers of 256 and 128 units. The first hidden layer is followed by sigmoid linear unit (
) activation, dropout with a rate of 0.1, and layer normalization; the second hidden layer uses
activation and dropout before the scalar output:
The reported implementation therefore incorporates matcher evidence through the selected acoustic-window embeddings and their matching logits. It does not use manually specified retrieval–verification posterior coefficients. The same top-five selection rule, verifier architecture, and aggregation setting are used across the reported random seeds and evaluation conditions.
Figure 2 summarizes the shared inference pathway and the Stage 4-only training signals. The inference model uses the structured query encoder, acoustic-window encoder, query–window matcher, top-five local verifier, and validation-selected operating threshold; the two teacher signals are absent at inference.
2.6. Controlled Ablation Instantiations of Tibetan-PASEM
The four stages are controlled instantiations of the same Tibetan-PASEM evidence-matching framework. Each instantiation enables one evidence source or training constraint while preserving the same Querybank198 data and evaluation protocol.
Stage 1, denoted PASEM-R, uses query-conditioned candidate retrieval only. It tests whether written-query evidence can retrieve plausible acoustic candidates.
Stage 2, denoted PASEM-RV, adds phonology-aware local verification. It tests whether top-ranked local acoustic evidence and query-side phonological features improve decision reliability under false-alarm pressure.
Stage 3, denoted PASEM-AS, introduces acoustically strengthened joint matching. It tests whether stronger acoustic-window evidence improves seen-query and confusable-negative discrimination.
Stage 4, denoted PASEM-TR, introduces teacher-guided transfer regularization during training. It is included to evaluate how auxiliary teacher information affects held-out written-query transfer under the same Querybank198 protocol.
The controlled instantiations measure how each evidence source changes performance on seen-query detection, unseen written-query transfer, and confusable-negative evaluation.
2.7. Teacher-Guided Transfer Regularization
Stage 4/PASEM-TR initializes the student from the selected Stage 3/PASEM-AS checkpoint and applies two fixed training-time signals. The first is the frozen PASEM-AS model defined in
Section 2.3,
Section 2.4 and
Section 2.5, including the structured written-query encoder, CNN-Conformer acoustic-window encoder, query–window matcher, and local verifier. It is fitted on seen_train, selected on seen_cv, contains approximately 4.68 million parameters, and is used with TTG = 2.0.
The second signal consists of archived pair-level probabilities generated by the MM-KWS-style reference implementation described in
Section 3.7 [
12]. The reference model is fitted on seen_train and selected on seen_cv. Its archived probabilities are used as fixed score-level targets with TMM = 1.0; its architecture, internal embeddings, and parameters are not incorporated into the Tibetan-PASEM student. The effective reference configuration, fitting schedule, checkpoint-selection rule, archive-status disclosure, and score-generation audit summary are reported in
Table S2 and File S1.
The Stage 4 objective is
where
is the supervised binary KWS loss,
transfers the softened posterior of the frozen PASEM-AS teacher,
matches the fixed MM-KWS-style probability targets,
preserves their pairwise score ordering, and
provides query-level contrastive regularization. The temperatures and loss coefficients are fixed across all reported Stage 4 runs [
12,
23].
Both teacher signals are frozen and used only during training. At inference, PASEM-TR uses only the trained Tibetan-PASEM student and the decision threshold defined by the evaluation protocol; neither teacher is loaded.
2.8. Operating-Point Selection and Metrics
For each configuration and random seed, the operating threshold τ* is selected on seen_cv by maximizing F1 subject to a false alarm rate (FAR) ≤ 0.10. Ties are resolved first by lower FAR and then by the higher threshold. The selected threshold is applied unchanged to the corresponding held-out evaluation conditions. Held-out labels, held-out score distributions, and test-specific thresholds are not used for operating-point selection [
10,
12,
15].
The complete stage-wise and query-view analyses use the original pair index. For the corrected-index re-evaluation,
is reselected independently for each random seed on cleaned seen_cv and is then applied unchanged to cleaned seen_test, unseen_test, and the confusable-negative condition. For the FAR–FRR visualization reported in
Section 3.4, the aligned seed-2028 score archives of PASEM-AS, PASEM-TR, and G-only/Direct-ST are evaluated over the same 27,577 cleaned seen_cv pairs and labels. The curves characterize the false alarm rate–false rejection rate (FAR–FRR) trade-off, and the marked points correspond to the thresholds selected by the same FAR-constrained validation rule. Held-out scores and labels are not used to construct the curves or select the marked operating points.
The reported thresholded metrics are Accuracy, Precision, Recall, F1, FAR, and FRR. Accuracy measures the overall pair-level decision rate. Precision measures the proportion of predicted triggers that are correct, whereas Recall measures the proportion of positive pairs detected by the system. F1 summarizes the balance between Precision and Recall. FAR measures false triggers among negative pairs, and FRR measures missed triggers among positive pairs. The area under the receiver operating characteristic curve (AUC) and equal error rate (EER) are additionally reported when aligned pair-level score archives are available; these two metrics summarize score-level discrimination across thresholds, whereas the remaining metrics describe performance at the validation-selected operating point.
Three-seed results are reported as the mean and sample standard deviation across seeds 2026, 2027, and 2028. Seed-level 95% confidence intervals for the principal unseen_test results are calculated as:
where
is the seed-level mean, s is the sample standard deviation, n = 3, and t
0.975,2 = 4.3027. These intervals characterize variation across the reported training seeds and should not be interpreted as pair-level binomial confidence intervals.
The split roles are fixed before evaluation. The seen_train split is used for model fitting, and seen_cv or cleaned seen_cv is used for model selection and operating-threshold selection under the corresponding protocol. The held-out conditions are used only for final metric computation. Test labels, held-out score distributions, and test-specific threshold searches are not used for model selection, query-view selection, teacher-signal construction, or operating-point selection.
2.9. Implementation Details and Protocol-Level Reproducibility
The experiments were developed within the open-source WeKWS keyword-spotting framework [
33]. WeKWS provides feature-extraction, checkpointing, and evaluation utilities, while the reported Tibetan-PASEM experiments add the structured written-query representation, Querybank198 pair construction, latent-window evidence matching, local verification, teacher-guided training settings, and split-aware evaluation workflow.
The reported system uses 16 kHz mono speech, 80-dimensional log-Mel filterbank features, a 25 ms frame length, a 10 ms frame shift, and approximately 2.0 s acoustic windows with a 0.8 s shift. The formal CNN-Conformer instantiation contains approximately 4.68 million parameters. The local verifier retains the five highest-scoring windows and uses a 256–128–1 multilayer perceptron with a dropout rate of 0.1. Additional model and training settings are specified in
Section 2.3,
Section 2.4,
Section 2.5,
Section 2.6,
Section 2.7 and
Section 2.8.
The released protocol-and-evaluation package contains the Querybank198 query metadata, original and cleaned pair indexes, negative-construction records, same-query control manifests, complete-file and decoded-PCM hash audits, acoustic-identity removal records, speaker and session-proxy statistics, development-to-held-out overlap matrices, threshold-selection and metric scripts, archived Tibetan-PASEM pair-level scores, per-seed thresholds, and score-alignment records [
24,
25,
26,
27].
These materials support reconstruction of the Querybank198 protocol, development-to-held-out audio correction, and score-level fixed-threshold evaluation. The package does not include the complete Tibetan-PASEM training implementation and therefore does not claim end-to-end model-training reproducibility. The accompanying README specifies the package scope, directory structure, required source resources, path configuration, expected split counts, score-file formats, evaluation commands, and known reproduction boundaries.