Published: 5 June 2026 · Direct Publication · arif-fazil.com/essays/
Epistemic Tag: CLAIM — awaiting adversarial review. Not peer-reviewed.
We identify a shared mathematical structure between seismic AVO anomaly analysis and transformer self-attention. Building on classical AVO formalisms (Rutherford & Williams; Castagna & Swan; Smith & Gidlow) and scaled dot-product attention (Vaswani et al.), we show that both systems compute a contrast between an observation and a context-specific baseline, followed by normalization that amplifies significant deviations. Under simplifying assumptions, we derive an explicit mapping between AVO residuals and attention logits, and propose governance frameworks—including GEOX ACRisk audits—for attention heads in geophysical AI models.
Published: 5 June 2026 · Direct Publication · arif-fazil.com/essays/
Epistemic Tag: CLAIM — awaiting adversarial review. Not peer-reviewed.
We identify a shared mathematical structure between seismic amplitude-versus-offset (AVO) anomaly analysis and transformer self-attention mechanisms used in modern machine learning models. Building on classical AVO formalisms (Rutherford & Williams, 1989; Castagna & Swan, 1997; Smith & Gidlow, 1987) and the scaled dot-product attention architecture (Vaswani et al., 2017), we show that both systems compute a contrast between an observation and a context-specific baseline, followed by a normalization step that amplifies significant deviations. In AVO, intercept–gradient analysis quantifies deviations from rock-physics trends such as the Mudrock line (Castagna et al., 1985), while in transformers, softmax-normalized query–key similarities yield a probability distribution that emphasizes tokens whose embeddings deviate most from the contextual norm. Under simplifying assumptions—linear background trend, approximately uniform attention baseline—we derive an explicit mapping between AVO residuals and attention logit residuals and demonstrate that both yield "winner-take-most" weighting of anomalous inputs. We discuss shared failure modes (false bright spots, spurious attention peaks), the role of priors and physical constraints, and implications for governance: attention heads in geophysical AI models can be audited using rock-physics-inspired anomaly metrics, including the GEOX ACRisk framework. This cross-disciplinary equivalence provides a conceptual and mathematical bridge between subsurface geophysics and neural sequence models, enabling transfer of interpretability and risk-governance tools across domains.
Two communities—exploration geophysicists and machine learning researchers—have independently developed formalisms for answering the same question: given a set of observations embedded in a context, which observations are salient, and by how much? The geophysicist answers it with intercept-gradient crossplots, background rock-physics trends, and fluid factors. The ML researcher answers it with query-key dot products, softmax normalization, and attention weight distributions. These communities rarely converse. Their tools share a mathematical skeleton they do not recognize in each other.
This paper names that skeleton.
The need is bidirectional. On the geophysics side, machine-learning-based seismic interpretation—particularly transformer architectures applied to pre-stack data—is proliferating (Li et al., 2025; seismic foundation models). Interpretability of these models is poor; a "bright spot" detected by a transformer attention head carries no physics guarantee that it corresponds to a Class III AVO anomaly rather than a tuning artifact. On the ML side, attention mechanisms lack external ground truth; there is no independent physical referent against which to validate that an attention peak is "real." The AVO community has spent four decades developing exactly such referents.
To our knowledge, no published work explicitly connects the mathematical structure of AVO anomaly detection to the softmax attention weight computation. The Pi-Transformer (Maleki & Pourmoazemi, 2025) introduces physics-informed priors for attention and is the closest bridge, but it does not reference geophysics or AVO. The connection identified in this paper appears to be novel.
CLAIM: Under specified conditions (linear background trend, uniform attention baseline, small-residual approximation), the AVO fluid factor ΔF = B − m·A − c and the attention logit residual δi = ei − ē play the same structural role in their respective systems. Both are contrast residuals. Both are normalized in ways that amplify outliers. Both produce "winner-take-most" outcomes. We do not claim that Shuey's approximation is identical to softmax, or that the Zoeppritz equations correspond to the QKV projection. The claim is structural: both systems instantiate a governed contrast amplification primitive over a context-specific baseline.
When a seismic P-wave encounters an interface between two rock layers with contrasting elastic properties, part of the energy is reflected. The reflection coefficient R depends on the angle of incidence θ. This angle-dependence—AVO—carries diagnostic information about pore fluids, lithology, and pressure (Ostrander, 1984).
Shuey (1985) gave the most widely used linearized form of the Zoeppritz equations:
where A is the intercept (normal-incidence P-wave reflectivity), B is the gradient (mid-offset AVO response), and C captures far-offset curvature. For angles below ~30°, the C term is often neglected, reducing the model to a two-parameter linear fit: R(θ) ≈ A + B sin²θ.
For any given reflection, one estimates A and B and plots the point (A, B) on a crossplot. Most reflections from brine-filled clastic rocks fall along an empirically calibrated background trend—the "Mudrock line" (Castagna et al., 1985). This trend represents the expected relationship Bbg(A) under the null hypothesis: no hydrocarbon, no anomalous pressure, no lithology change beyond the background.
A hydrocarbon-charged reservoir produces a reflection whose (A, B) coordinates deviate from this background. The deviation:
is the AVO anomaly measure. Large |ΔB| → anomalous reflection.
Smith & Gidlow (1987) introduced the "fluid factor" ΔF, which projects (A, B) onto a direction orthogonal to the brine-sand trend in crossplot space. For a linear background B = mA + c:
The fluid factor is a contrast score—a scalar whose magnitude indicates how far a given observation deviates from the null (brine) hypothesis. It is not a binary flag; it is a continuous measure of anomalousness.
Rutherford & Williams (1989) and Castagna & Swan (1997) classified gas-sand reflections into four classes (I–IV) based on their position on the A-B crossplot relative to the background trend. This classification is a discretization of the continuous contrast measure ΔB. Class III (bright spots, strongly negative A with increasing magnitude at far offsets) exemplifies the classical anomaly. Class IV (decreasing amplitude with offset) is a subtler, often-missed signal—the false-negative case.
Vaswani et al. (2017) defined scaled dot-product attention as:
For a single query vector q and a set of key-value pairs {kj, vj}j=1N:
where the attention weights αj are computed via softmax over alignment scores:
The softmax has two critical effects on the raw alignment scores {ej}:
Consider a simplified scenario: one key ki has alignment ei = ē + δ where ē is the baseline score shared by all other keys, while the remaining N−1 keys have score ej≠i = ē. Then:
If δ = 0 (no contrast), αi = 1/N (uniform baseline). If δ > 0, αi > 1/N. For moderate N and growing δ, αi → 1. The first-order Taylor expansion around δ = 0:
shows that attention weight is, to first order, linear in the contrast δ, with exponential higher-order terms that accelerate convergence to 1 for larger deviations. This is the mathematical basis for the "winner-take-most" characterization.
When all ej are equal, attention defaults to a uniform distribution over keys. This uniform distribution serves as an implicit baseline prior—the "no-salient-token" null hypothesis. When a token's alignment exceeds this baseline, the softmax rejects the null and allocates probability mass disproportionately to that token. The baseline is not explicitly modeled; it is a consequence of the softmax functional form.
We now formalize the interpretation of attention weights as governed contrast measures. This section establishes the precise mathematical vocabulary used in the equivalence mapping (Section 4).
For a fixed query q and a set of N keys {kj}, define the alignment scores ej = q·kj / √dk. Let the baseline alignment be the arithmetic mean:
Then the contrast residual for key ki is:
This residual measures how much key ki deviates from the average compatibility with q. It is the attention-domain analog of the AVO fluid factor ΔF.
The attention weight αi can be expressed directly in terms of δi:
This makes explicit that the baseline ē cancels—only contrasts matter. The attention distribution is determined entirely by the pattern of deviations from the mean, not by the absolute magnitude of the alignment scores.
These properties establish attention as a contrast-governed mechanism. The baseline—equal relevance—is the implicit null hypothesis. The softmax is the amplifier. The weight αi is the amplified contrast response.
Table 1 presents the side-by-side decomposition of AVO anomaly detection and transformer self-attention into six shared computational stages.
Table 1: Stage-by-stage structural mapping between AVO anomaly detection and transformer self-attention.
| Stage | Seismic AVO Anomaly Detection | Transformer Self-Attention |
|---|---|---|
| 1. Observation | Reflection coefficient R(θ) as function of incidence angle (pre-stack amplitude data) | Sequence of token vectors x1, …, xN (word embeddings, time series, seismic trace segments) |
| 2. Feature Extraction | Intercept A and gradient B from R(θ) via Shuey's two-term linearization (Shuey, 1985) | Query q = WQx and Keys kj = WKxj from learned linear projections (Vaswani et al., 2017) |
| 3. Baseline | Empirical background trend Bbg(A) — e.g., Mudrock line for brine-saturated clastics (Castagna et al., 1985) | Implicit uniform prior — all keys equally relevant (no salient token). Emergent property of softmax when all ej equal. |
| 4. Contrast Computation | Residual ΔB = Bobs − Bbg(Aobs), or fluid factor ΔF = B − mA − c (Smith & Gidlow, 1987) | Contrast residual δi = ei − ē; alignment scores ej = q·kj / √dk |
| 5. Normalization & Amplification | Scale by background variance σbg; use Mahalanobis distance in (A, B) space; threshold at kσ for anomaly flag | Softmax: αj = exp(ej) / Σℓ exp(eℓ). Normalizes to distribution + exponentially amplifies differences |
| 6. Outcome | Anomaly indicator: yes/no, AVO Class I–IV (Rutherford & Williams, 1989), continuous anomaly strength | Attention weight distribution {αj} — selects salient token(s), down-weights rest. "Winner-take-most." |
Each row in Table 1 identifies a functional correspondence. The operations are not identical in implementation—linear regression vs. learned projections, linear subtraction vs. softmax exponential—but they are identical in structural role. Both systems: observe, extract features, establish a baseline, compute contrast residuals, normalize/amplify those residuals, and produce a "spotlight" outcome.
Assume the AVO background trend is linear: Bbg(A) = mA. (The extension to B = mA + c is straightforward; we set c = 0 without loss of generality by translating coordinates.) For a specific reflection with intercept A0 and observed gradient B0, define:
The anomaly residual vector is:
Let u⊥ be a unit vector perpendicular to the background line in (A, B) space. The anomaly magnitude—the projection of the residual onto the direction of maximum fluid sensitivity—is:
This is the fluid factor ΔF of Smith & Gidlow (1987), expressed in the coordinate system where u⊥ is the fluid-sensitive direction. The anomaly magnitude is a signed scalar contrast residual.
Fix a query q. Let {ej}j=1N be the alignment scores. Define the baseline logit:
The contrast residual for key ki is δi = ei − ē. For the special case where exactly one key deviates from the baseline (the "one-outlier model"), with δi = δ and δj≠i = 0, the attention weight is:
Key properties of this function:
Linearizing the softmax response around δ = 0:
The zero-order term (1/N) is the baseline—uniform attention. The first-order term ((N−1)/N²)·δ is the contrast response—the attention weight's sensitivity to the residual. Comparing to the AVO case:
In both systems, the response is, to first order, a monotone function of a contrast residual. The AVO fluid factor ΔB⊥ and the attention logit residual δ play the same formal role: they are the input to a normalization-amplification stage that produces the output—anomaly flag in AVO, attention weight in transformers.
The formal mapping holds under the following conditions:
Outside these conditions, the mapping is conceptual rather than formal: both systems still compute contrast + normalize + amplify, but the functional forms diverge.
AVO domain: False positive anomalies—"false bright spots"—occur when the background trend is misestimated or when non-hydrocarbon effects (shallow high-porosity brine sands, thin-bed tuning, processing artifacts) mimic a fluid response. Castagna et al. (1985) documented that non-hydrocarbon reflections can exhibit increasing AVO, producing false positives when the background model is inadequately conditioned for the local stratigraphy (Ostrander, 1984; Castagna & Swan, 1997).
Attention domain: When many keys are similar or the query is ambiguous, attention may focus on an irrelevant token—the "attention misfire" (Jain & Wallace, 2019; Wiegreffe & Pinter, 2019). This occurs when the learned query-key similarity is biased or when out-of-distribution inputs confuse the scoring mechanism. The attention head "detects" a contrast that has no task-relevant significance.
Structural symmetry: Both failures stem from baseline misspecification. The contrast amplification mechanism faithfully amplifies whatever deviates from the assumed baseline—if the baseline is wrong, the amplification is misleading. A false bright spot is a spurious attention peak in (A, B)-space; a spurious attention peak is a false bright spot in embedding-space.
AVO domain: Class IV gas sands decrease in amplitude with offset (Rutherford & Williams, 1989; Castagna & Swan, 1997), contradicting the classical bright-spot model. They can be missed entirely if the anomaly detector is tuned only for amplitude increases, producing a false negative.
Attention domain: A truly important token may fail to receive high attention weight if its significance is masked by a noisy context—many moderately relevant tokens dilute the contrast distribution, and no single token stands out (the "diffuse attention" regime).
Structural symmetry: Both failures occur when genuine signal does not produce sufficient contrast against the prevailing context. The signal is present but does not "break the pattern" strongly enough to survive normalization.
The GEOX ACRisk framework (Arif, 2026; GEOX SPACE KNOWLEDGE ARTIFACT) provides a structured audit for anomalies detected by geophysical AI models. The framework decomposes any detected anomaly into three components:
We propose extending ACRisk to attention heads in geophysical transformer models. For each attention head that "flags" a seismic feature as salient (high αj):
This framework gives teeth to the formal equivalence: it is not merely an intellectual curiosity. It provides a governance mechanism—a way to audit whether an AI-detected "anomaly" in seismic data is grounded in the same physics that a human interpreter would demand. The ACRisk audit turns the mathematical parallel into an operational quality gate.
Both AVO and attention operate under non-stationary conditions (Envisioning, 2025). In AVO, the background trend shifts with depth, compaction, diagenesis, and basin history—the Mudrock line calibrated in the Gulf of Mexico does not transfer to the Malay Basin without adjustment. In transformers, the "background" distribution of token relationships shifts with domain, language, and temporal drift in training data. Both systems must function under drifting priors—the baseline is never static. Governance frameworks (ACRisk, F13 sovereign review) must account for this: an anomaly that is genuine in one context may be baseline in another.
The foundational AVO literature spans Ostrander (1984), who first demonstrated gas-sand AVO effects; Shuey (1985), who provided the linearized approximation; Rutherford & Williams (1989), who introduced the four-class AVO system; Smith & Gidlow (1987), who introduced the fluid factor as a contrast measure orthogonal to the brine trend; and Castagna & Swan (1997), who systematized crossplot interpretation. The Castagna Mudrock line (Castagna et al., 1985) remains the most cited empirical background trend. This body of work established the conceptual framework of deviation from a calibrated background as the signal of interest.
Vaswani et al. (2017) introduced scaled dot-product attention. Subsequent work has analyzed attention through multiple lenses: as content-based retrieval (Graves et al., 2014), as a probabilistic inference mechanism, and as an interpretability challenge (Jain & Wallace, 2019; Wiegreffe & Pinter, 2019). The "winner-take-most" property of softmax attention is well-characterized in the ML literature but has not been connected to geophysical contrast detection.
The Pi-Transformer (Maleki & Pourmoazemi, 2025) introduces a physics-informed prior distribution over attention weights and treats deviations from this prior as anomalies. This is architecturally analogous to using the Mudrock line as a prior for AVO analysis—both introduce domain knowledge to constrain what counts as "anomalous." However, Maleki & Pourmoazemi do not reference geophysics or AVO. The connection identified here has not been drawn elsewhere.
The broader literature on contrastive methods—contrastive learning (Chen et al., 2020), contrastive predictive coding (van den Oord et al., 2018), and anomaly detection via contrast—suggests that learning by distinguishing signal from background is a deep organizational principle across machine learning. This paper adds to that literature by showing that the principle extends to classical geophysical signal processing, unifying a four-decade-old exploration technique with the core mechanism of modern AI architectures.
CLAIM: To the best of our knowledge, no prior publication has drawn an explicit formal parallel between AVO intercept-gradient analysis and transformer self-attention. The closest work (Pi-Transformer, physics-informed attention) converges on the same pattern without naming the AVO connection. This paper fills that gap.
The attention mechanism's probabilistic formulation suggests treating the AVO background trend not as a deterministic line but as a probability distribution. Instead of a binary anomaly flag, one could compute a posterior probability that a given (A, B) point belongs to the background (brine) population versus an anomalous (hydrocarbon) population—a Bayesian fluid factor using softmax-like normalization over competing rock-physics hypotheses. This would provide geophysicists with a graded confidence measure rather than a hard threshold.
Transformers use multiple attention heads, each learning a different projection. By structural analogy, one could design a "multi-head AVO" framework where different rock-physics models (Mudrock line, Gardner trend, local basin calibration, depth-dependent trends) serve as parallel "heads," with each head's anomaly score combined through a learned or physics-informed gating mechanism. This would generalize the single-trend AVO analysis while preserving physical interpretability.
Rather than relying on hand-calibrated background trends that are basin-specific and depth-dependent, one could train a model to learn a context-dependent baseline from data—a transformer that learns what "normal" reflectivity looks like in a given geological setting and flags deviations. This is the geophysical analog of the Pi-Transformer, with the additional benefit that the learned baseline can be inspected against known rock-physics relationships.
AVO interpreters already use crossplots as interpretability tools. Attention weight visualization—showing which parts of the seismic trace contributed most to an anomaly classification—could provide a new dimension of interpretability for machine-learning-based seismic interpretation, grounding AI decisions in physically inspectable amplitude-versus-angle behavior. A transformer that classifies an event as a Class III bright spot could be asked to show why, in the form of the attention weights over angle-gather traces, which can then be cross-checked against the expected AVO signature.
At the highest level, this equivalence suggests that "contrast detection and amplification" is a computational primitive—a reusable, domain-independent operation that can be instantiated in different substrates (physics-based regression for AVO, learned projections for attention) but always performs the same information-processing function: find what deviates from context, then amplify its weight. Identifying this primitive opens the door to formal comparisons between other apparently unrelated domains—any system that computes "what's different" and "how much it matters" may be an instance of the same kernel.
Seismic AVO anomaly detection and transformer self-attention both instantiate a contrast-governed anomaly detection primitive: establish a context-specific baseline, compute contrast residuals, normalize and amplify those residuals, and produce a spotlighted outcome. In AVO, the contrast is between observed and expected reflection behavior via intercept-gradient trends; in attention, it is between a query-key similarity and the distribution of similarities across the context window.
Under idealized conditions—linear background trend, uniform attention baseline, small-residual approximation—we have shown that the AVO fluid factor ΔF and the attention logit residual δ play the same structural role, and that the softmax attention weight αi is an amplified normalization of δ analogous to thresholding ΔF for anomaly classification. The correspondence is not one-to-one in functional form (linear subtraction vs. softmax exponential) but is exact in computational architecture: both systems compute "what's different, then turn up its weight."
This equivalence is actionable. The GEOX ACRisk framework—originally designed for auditing human-interpreted seismic anomalies—extends naturally to auditing attention-head outputs in geophysical transformers. An attention peak without a corresponding AVO contrast is a false bright spot in embedding-space. A genuine AVO anomaly that fails to produce an attention peak is a false negative—a Class IV sand that the model missed. The audit framework gives governance teeth to the mathematical parallel.
The formal insight that governed contrast extraction is at play in both domains opens a dialog between geophysicists and AI researchers. Future work includes: empirical evaluation of ACRisk-audited attention heads on pre-stack seismic transformers; development of multi-head AVO frameworks using parallel rock-physics priors; and generalization of the contrast primitive to other domains where "find what stands out from context" is the core computational task.
We are all, as it turns out, in the business of finding what stands out. The difference is only in what we count as background.
By Arif Fazil Sealed 999 · 23 min read
Muhammad Arif bin Fazil
Geoscientist · Architect, arifOS · Petronas Carigali · UW–Madison '13
Penang, Malaysia
Published: 05 June 2026 · Direct Publication · /words/ context
Epistemic Tag: INT — interpretive synthesis across AI governance, institutional economics, and systems theory
Pairs with: Physics-Constrained Attention: Zoeppritz (EUREKA II) · The Contrast Primitive Derivation (EUREKA III)