4.1. F3 Data and Experimental Protocol
The F3 dataset comprises public marine three-dimensional seismic data and well information from the offshore of the Netherlands [
7]. We chose F3 because its seismic neighborhoods, well-log windows, and location/interval metadata can be aligned within one survey, providing a reproducible setting for seismic–log proxy ranking. The local paired seismic–well index contains 364 rows. After filtering for valid proxy scores, 321 samples remain. The four contributing wells are F03-2 with 91 valid samples, F03-4 with 82, F06-1 with 80, and F02-1 with 68. Well inclusion is determined by the availability of valid paired seismic, log, and target components. Each sample links a seismic neighborhood to a well-log window and metadata describing its well, inline/crossline location, time interval, and horizon interval. Our main evaluation uses sample-stratified splits, so intervals from the same well can appear in both training and test sets.
Figure 2 distinguishes the four well coordinates from the multiple scored intervals associated with each coordinate.
The model ranks paired intervals with available seismic context and log windows. The spatial evaluation is concentrated at four well coordinates within F3.
Fluid migration interpretation is motivated by several types of evidence. The seismic image provides reflection geometry and local structural texture. Chimney response highlights vertically disturbed zones that can be associated with migration pathways or gas-related effects. Fault enhancement highlights discontinuities or fault-like features that can provide migration conduits. Well logs provide local petrophysical and elastic responses near the borehole. We visualize this context in
Figure 3.
The three panels in
Figure 3 illustrate geological context for fluid migration assessment. The seismic image contains stratigraphic and structural texture, while the attribute views and logs provide complementary context. The ranking experiment uses paired seismic and log inputs to predict the favorability proxy defined below.
Each paired sample has a continuous fluid-migration favorability proxy score . The target-generation script ranks available chimney response, fault enhancement, seismic discontinuity, and log-anomaly measures within horizon intervals and combines them with a horizon-position indicator at weights 0.35, 0.30, 0.20, 0.10, and 0.05, respectively; weights for missing components are renormalized. Valid samples require available chimney and fault proxies. The score is an attribute-based index of relative screening priority, with larger values denoting higher favorability. Its components draw on seismic and log information related to the model inputs. Proxy ranks are computed on the paired index before the downstream train/test split, and the benchmark evaluates recovery of this predefined score.
Figure 4 shows the chimney and fault responses used to provide geological context for this target.
The displayed responses indicate structural setting and complement the seismic and well-log evidence used for ranking.
For the auxiliary binary task, the high-favorability cutoff is the upper quartile of the continuous scores in each seed’s training set
, as shown in Equation (
13):
In Equation (
13),
is the indicator function. The resulting training-score thresholds are 0.68854, 0.69392, and 0.69559 for seeds 0, 1, and 2, respectively; each seed-specific threshold is then applied unchanged to its validation and test samples. The continuous score remains the primary target.
Table 1 summarizes the sample counts and target distribution for the benchmark.
Figure 5 contrasts one high-favorability and one low-favorability paired sample.
The examples have different local seismic textures and well-log responses, but neither modality alone fully defines the target. This motivates the architecture in
Section 3: the SFM branch captures seismic texture and structural context, the well-log token encoder captures borehole response, and the fusion module combines them for ranking.
The experiments evaluate the framework in
Section 3 as a controlled ranking problem. Each model receives samples from the same F3 paired seismic and well-log benchmark, and the primary target is the continuous fluid-migration favorability score described in this section. The main question is whether a model can rank paired samples so that higher-favorability candidates appear near the top of a review queue.
We compare nine model configurations that vary the input modality, seismic representation, and fusion module. The log-only model removes the seismic branch. The seismic-only CNN removes the well-log token encoder. CNN concatenation, cross attention, and learned query fusion use the task-trained CNN with the corresponding modules in Equations (7)–(9). Feature-only SFM uses the frozen tokens and their trainable projection without logs. The three SFM multimodal configurations combine the same frozen tokens with concatenation, cross attention, or learned query fusion.
Table 2 reports the trainable parameter counts obtained from the saved downstream checkpoints.
The SFM Base 224 checkpoint contains 88,878,336 parameters, of which 85,405,440 belong to the encoder used for feature extraction; the masked-autoencoder decoder is unused downstream. All of these SFM parameters remain frozen. The task-trained CNN seismic encoder contains 125,731 parameters. The comparison therefore contrasts a large pretrained, frozen encoder with a smaller encoder learned from the downstream data.
For each seed, the 321 valid samples are stratified by ranked target quantiles and divided into 224 training samples (69.8%), 48 validation samples (15.0%), and 49 test samples (15.3%). Seeds 0, 1, and 2 determine both this split and model initialization. The continuous target is standardized with the corresponding training-set mean and standard deviation. Every model is trained for 50 epochs with batch size 16 using AdamW, a learning rate of , weight decay of , and a gradient-norm limit of 1.0; the reported checkpoint is the final epoch. Frozen SFM representations are precomputed, whereas all CNN weights are learned from the 224 downstream training samples. The models share downstream splits and optimization budgets, with the SFM additionally drawing on its pretraining corpus. The runs were executed on an NVIDIA RTX 3090 Ti GPU.
The two-layer well-log token encoder has 1.19 million trainable parameters, and the multimodal models have 1.35–1.40 million trainable downstream parameters. To measure sensitivity to training-set size, we train the log-only and seismic-only CNN configurations on nested fractions of each training split while holding validation and test samples fixed.
Table 3 reports the resulting learning curves for the continuous task.
The learning curves show little change between the two largest training fractions under sample-stratified evaluation. In paired comparisons, moving from 75% to 100% of the training split changes log-only nDCG@20 by percentage points (95% CI [, 11.67]) and seismic-only CNN nDCG@20 by points (95% CI [, 8.59]). The corresponding intervals for Spearman and RMSE also include zero: log-only changes are 1.09 points [, 54.95] and 0.14 points [, 3.95], while CNN changes are points [, 0.22] and 0.25 points [, 0.68]. Across the three seeds, the added quarter of training samples produces no detectable gain in these metrics.
We further test whether the log-only result depends on the 2048-dimensional feed-forward sublayer.
Table 4 reduces that width to 512 and 256 while retaining two Transformer layers, four attention heads, the input projection, and the original optimization protocol.
Relative to the 2048-width encoder, the paired nDCG@20 differences are points (95% CI [, 4.67]) for width 256 and points (95% CI [, 6.61]) for width 512. The observed means are similar across widths, while the confidence intervals allow both gains and losses.
Neighboring windows overlap, and sample-stratified splits can share well-specific structure between training and test sets. We therefore assess performance at unseen wells through leave-one-well-out evaluation. Each fold holds out one complete well with 68–91 samples; the remaining wells provide 184–203 training and 46–50 validation samples. For each seed, metrics are averaged over the four held-out wells before the three-seed confidence interval is computed.
Table 5 reports this well-macro result.
Compared with sample-stratified evaluation, well-macro nDCG@20 decreases by 8.39 points for log only, 6.16 points for the seismic-only CNN, and 8.91 points for SFM feature only. These decreases quantify the gap between sample-stratified and unseen-well evaluation for the three single-modality models. The training-size and capacity analyses characterize sensitivity within the sampled wells, while the held-out results reveal the additional challenge of generalizing across well locations.
A positional-encoding sensitivity run also tests the implementation choice in
Section 3.3. Adding fixed sinusoidal position embeddings to the otherwise unchanged log-only model yields 81.25% nDCG@20, 27.04% Spearman, and 18.22% RMSE, compared with 81.20%, 27.01%, and 17.65%, respectively, without positional embeddings. nDCG@20 and Spearman change little, while RMSE increases, giving no consistent improvement from the added encoding in this experiment.
nDCG@20 is the primary metric because the intended output is a ranked candidate list. We also report nDCG@10 as a complementary shorter cutoff because a 20-sample queue is a substantial fraction of a 49-sample test split. Spearman correlation measures monotone agreement between predicted and target scores. Root mean square error (RMSE) measures absolute score error, and measures explained variance relative to the test set target distribution. The auxiliary binary results use area under the receiver operating characteristic curve (AUROC), area under the precision–recall curve (AUPRC), and balanced accuracy. For readability, all bounded scores and error values are reported as percentages, with RMSE scaled to the normalized 0 to 1 target.
4.2. Main Ranking Results
Figure 6 summarizes the eight core configurations, while
Table 6 also includes the matched CNN learned-query control added for the backbone comparison.
Table 6 reports that the frozen SFM configuration with learned query fusion has the highest observed nDCG@20 mean among the listed configurations. Its reported nDCG@20 is 85.15%, with Spearman correlation of 41.29%, RMSE of 15.60%, and
of 17.21%. These measures describe complementary aspects of the continuous task. nDCG@20 evaluates the head of the ranked queue, Spearman evaluates monotone agreement, RMSE evaluates score error, and
evaluates explained variation relative to a mean predictor.
The learned query fusion module maps the combined seismic and log token sequences to a compact sample-level representation for the continuous ranking task.
The first experimental pattern concerns multimodal input. The strongest listed single-modality model uses only well logs and reports nDCG@20 of 81.20%, whereas the frozen SFM configuration with learned query fusion reports 85.15%. By Equation (
12), the reported difference is
. This is the largest listed comparison for the main ranking metric and suggests that combining seismic context with borehole response can improve fluid migration ranking over either modality alone in the evaluated F3 setting.
The second pattern compares the leading frozen SFM and CNN configurations. CNN concatenation reaches nDCG@20 of 83.84%, whereas frozen SFM features with learned query fusion reach 85.15%. The displayed values differ by , reflecting the combined contribution of the seismic representation and fusion configuration. A matched-fusion comparison holds learned query fusion fixed: the CNN version reaches 83.36%, giving an SFM-minus-CNN difference of 1.80 points (paired 95% CI [, 4.80]). The mean favors the SFM backbone, with uncertainty spanning zero across the three seeds.
nDCG@10 provides a complementary view of the short ranked queue. CNN concatenation reaches 83.44%, whereas frozen SFM features with learned query fusion reach 80.40%. Together with nDCG@20, these results characterize ranking behavior at different review depths.
The final ranking observation concerns calibration versus ordering. SFM feature-only has lower nDCG@20 than log-only, yet its RMSE is slightly better than the log-only RMSE. This pattern shows that score accuracy and prioritized ranking provide complementary views of model performance. For a screening workflow, ranking behavior directly reflects the ordering of high-favorability samples in the review queue.
4.4. Fusion Strategy and Qualitative Analysis
The frozen SFM variants compare fusion modules while holding the seismic feature source fixed. SFM concatenation reports 82.13% nDCG@20, SFM cross attention reports 83.83%, and SFM learned query fusion reports 85.15%. In this benchmark, the learned-query configuration has the highest observed nDCG@20 mean among the listed frozen-SFM variants.
The frozen SFM encoder produces a token sequence that is summarized by the downstream module. Pooled concatenation, cross attention, and learned query fusion provide progressively different forms of interaction and aggregation. The present results characterize their empirical ranking behavior.
The nDCG@20 confidence intervals in
Table 6 overlap across fusion variants, leaving their mean ordering uncertain across splits and seeds.
At nDCG@10, SFM cross attention reaches 81.69%, SFM learned query fusion reaches 80.40%, and SFM concatenation reaches 80.32%. The learned query result is therefore specific to the nDCG@20 evaluation used as the primary metric.
Figure 8 relates representative learned-query predictions to their seismic and well-log inputs.
The illustrated high-score cases connect the seismic structure and log response to the constructed favorability proxy. Chimney and fault views show the corresponding target components.
Some high-favorability samples receive moderate predictions and show mixed seismic and log cues.
The qualitative results also distinguish the ranking and binary tasks. Ranking rewards a model for keeping several intervals with high fluid migration favorability in the upper part of the list. The binary task instead applies the seed-specific training-set threshold in Equation (
13). This makes ranking the more direct evaluation for the intended screening workflow.