Published: 5 June 2026 · Direct Publication · arif-fazil.com/essays/
Trilogy: Essay #11 (The Bridge) → Essay #12 (The Consequences) → Essay #13 (The Derivation)
Epistemic Tag: CLAIM — awaiting adversarial review. Not peer-reviewed.
The third and final paper in the EUREKA trilogy. We present the explicit, step-by-step mathematical derivation establishing the structural equivalence between the AVO fluid factor ΔF and the transformer attention logit residual δᵢ. Starting from the exact Zoeppritz equations through the Aki–Richards linearization and Shuey two-term approximation, and in parallel from the scaled dot-product attention mechanism through softmax normalization, we derive both contrast operators in a shared notational framework. We provide a first-order Taylor expansion proving that both systems compute a residual against a calibrated baseline, then amplify deviations. Every condition, every assumption, and every boundary of equivalence is explicitly stated. This paper is the mathematical lock-in: the skeleton, bone by bone, with every joint labeled.
Published: 5 June 2026 · Direct Publication · arif-fazil.com/essays/
Trilogy: Essay #11 (The Bridge) → Essay #12 (The Consequences) → Essay #13 (The Derivation)
Epistemic Tag: CLAIM — awaiting adversarial review. Not peer-reviewed.
This paper presents the explicit, step-by-step mathematical derivation establishing the structural equivalence between the seismic AVO fluid factor ΔF (Smith & Gidlow, 1987; Fatti et al., 1994) and the transformer attention logit residual δi (Vaswani et al., 2017). We trace both formalisms from their exact foundations—the Zoeppritz equations for AVO, the scaled dot-product attention for transformers—through their respective linearizations and approximations, to their operational forms. In a shared notational framework, we derive both contrast operators and prove via first-order Taylor expansion that each implements a three-stage primitive: extract features, compute residual against a calibrated baseline, amplify deviations. The AVO residual ΔB = Bobs − (mAobs + c) and the attention residual δi = ei − ē are shown to play identical structural roles. We explicitly state every condition under which the mapping holds and every boundary at which it breaks. This paper is the mathematical lock-in for the EUREKA announced in Essay #11 and the governance framework proposed in Essay #12.
| Essay | Title | Function | Question Answered |
|---|---|---|---|
| #11 | Contrast-Governed Anomaly Detection: A Formal Bridge | The EUREKA — structural mapping | What is the connection? |
| #12 | Physics-Constrained Attention: Zoeppritz as Constitutional Floor | The consequences — governance, hallucination prevention | So what? What must we build? |
| #13 | The Contrast Primitive Derivation (this paper) | The mathematical lock-in — bone-by-bone proof | How exactly? Prove it. |
Table 0: The trilogy architecture. This paper (#13) supplies the rigorous derivation that #11 gestured toward and #12 built upon.
This paper is self-contained. The reader need not have read Essays #11 or #12. All notation is defined from first principles. All derivations proceed step by step. Every equation carries its assumptions explicitly.
| R(θ) | P-wave reflection coefficient at incidence angle θ |
| Vp, Vs, ρ | Compressional velocity, shear velocity, bulk density |
| ΔVp/Vp, etc. | Fractional elastic contrast across the interface |
| A | AVO intercept — normal-incidence reflectivity R(0) |
| B | AVO gradient — coefficient of sin²θ in Shuey two-term |
| m, c | Mudrock line parameters: Bbg = mA + c (Castagna et al., 1985) |
| ΔB | Contrast residual: Bobs − Bbg(Aobs) |
| ΔF | Fluid factor: weighted combination of Rp and Rs (Smith & Gidlow, 1987) |
| q, kj | Query vector, j-th key vector ∈ ℝdk |
| N | Number of keys |
| ej | Alignment logit: ej = q·kj / √dk |
| ē | Mean logit: ē = (1/N) Σj ej |
| δj | Logit residual: δj = ej − ē |
| αj | Attention weight: αj = softmax(ej) |
When a plane P-wave impinges on a planar elastic interface at angle θ, four wave modes are generated: reflected P, reflected SV, transmitted P, transmitted SV. Conservation of energy and continuity of displacement and stress across the interface yield a 4×4 linear system:
M(θ, Vp1, Vs1, ρ1, Vp2, Vs2, ρ2) · R = b(θ)
where R = [RP, RS, TP, TS]T contains the four unknown reflection and transmission coefficients. The matrix M is dense with trigonometric functions of all four wave angles and elastic property ratios. The solution is exact but algebraically opaque: RP(θ) is a rational function of six elastic parameters and θ, with no separation of Vp, Vs, and ρ contributions.
Assumption 1 (Small contrasts): |ΔVp/Vp| ≪ 1, |ΔVs/Vs| ≪ 1, |Δρ/ρ| ≪ 1.
Under Assumption 1, the Zoeppritz system linearizes to first order in the fractional contrasts. The Aki–Richards approximation expresses R(θ) as a weighted sum of three independent contrast terms:
R(θ) ≈ ½(ΔVp/Vp + Δρ/ρ)(1 + tan²θ) − 2(Vs/Vp)²(ΔVs/Vs + ½Δρ/ρ) sin²θ + ½(Δρ/ρ)(1 − 4(Vs/Vp)² sin²θ)
Term 1: P-velocity + density contrast (dominates near-normal). Term 2: S-velocity contrast (dominates at mid-angles). Term 3: Density contrast (small at moderate angles).
Shuey substituted Poisson's ratio σ for Vs and reorganized Aki–Richards into angle-separated terms:
R(θ) ≈ A + B sin²θ + C sin²θ tan²θ
where:
A = R(0) = ½(ΔVp/Vp + Δρ/ρ)
B = ½(ΔVp/Vp) − 4(Vs/Vp)²(ΔVs/Vs) − 2(Vs/Vp)²(Δρ/ρ)
C = ½(ΔVp/Vp)(tan²θ − sin²θ) / (sin²θ tan²θ) ≈ ½(ΔVp/Vp) at wide angles
Assumption 2 (Moderate angles): θ ≲ 30°. Then sin²θ tan²θ ≪ sin²θ and the C-term is negligible.
Under Assumption 2, the two-term Shuey approximation is the workhorse:
R(θ) ≈ A + B sin²θ
For brine-saturated clastic rocks, Vp and Vs follow an empirical linear relationship. This propagates through Shuey's coefficients to yield a linear relationship between A and B for non-anomalous (brine-saturated) reflectors:
Bbg(A) = m·A + c
where m (negative) and c are calibrated from well-log data or regional rock-physics trends. For the Wiggins et al. (1983) constant Vp/Vs = 2.0 case, m = −1 under the approximation B ≈ A − 2Rs.
For an observed reflector with extracted attributes (Aobs, Bobs):
ΔB = Bobs − Bbg(Aobs) = Bobs − m·Aobs − c
Property 1: If the reflector is brine-saturated and on the Mudrock line, ΔB ≈ 0.
Property 2: If the reflector contains hydrocarbons, Poisson's ratio drops, Vp/Vs decreases, and |ΔB| > 0.
The operation is: subtract the expected background, keep the deviation.
The fluid factor re-projects the (A, B) or (Rp, Rs) pair onto a "fluid-sensitive" axis orthogonal to the brine trend:
ΔF = Rp − (1.16 / γ) Rs
where γ = Vp/Vs (background) and 1.16 is the Mudrock slope. By construction: ΔF ≈ 0 for brine; |ΔF| > 0 for gas. The fluid factor is a specific linear combination of the same residual ΔB expressed in reflectivity coordinates. Both share the identical structure:
Residual = Observation − Calibrated_Baseline(Observation)
Rutherford & Williams (1989) and Castagna & Swan (1997) classify reflectors on the A–B plane:
| Class | A (Intercept) | B (Gradient) | ΔB |
|---|---|---|---|
| I | High + | Negative | Below Mudrock line |
| II | ~Zero | Negative | Below Mudrock line |
| III | High − | Negative | Far below Mudrock line |
| IV | High − | Positive | Above Mudrock line (counter-intuitive) |
The classification rule is a threshold on ΔB: Anomaly = Θ(|ΔB| − τ) where Θ is a step function and τ is a calibrated cutoff.
Stage 1 — Extract: (Aobs, Bobs) from pre-stack gathers via Shuey two-term fit.
Stage 2 — Residual: ΔB = Bobs − (m·Aobs + c). Zero on background, non-zero off.
Stage 3 — Amplify/Classify: |ΔB| > τ → anomaly flagged (Class I–IV).
The amplification function Θ is linear thresholding. The signal is the residual.
Given a query q ∈ ℝdk and N keys {k1, …, kN} ⊂ ℝdk, the scaled dot-product attention computes:
Attention(q, K, V) = Σj=1N αj vj
where the attention weights are:
αj = softmax(ej) = exp(ej) / Σi=1N exp(ei)
with alignment logits:
ej = (q · kj) / √dk
The √dk scaling prevents the dot product from growing with dimension and pushing softmax into saturated (near-one-hot) regimes. This is the full, exact attention computation — O(N²) in key count, with an exponential nonlinearity.
The softmax function has two essential properties relevant to our mapping:
Property 1 (Normalization): Σj αj = 1. The output is a valid probability distribution.
Property 2 (Exponential amplification): ∂αi/∂ei = αi(1 − αi). The gradient is maximal at αi = 0.5 and vanishes as αi → 0 or 1. This means softmax disproportionately amplifies differences between the largest logits while suppressing small ones.
The softmax is not merely a normalization — it is a contrast amplifier. Given any variation in {ej}, it concentrates probability mass onto the maximum. The larger the spread of logits, the more extreme the concentration.
Definition (Uniform Baseline): When all keys are equally relevant to the query, the logits are identical:
e1 = e2 = … = eN = E (constant)
Then ē = E, all residuals δj = 0, and:
αj = exp(E) / (N·exp(E)) = 1/N for all j
The uniform distribution αj = 1/N is the attention Mudrock line: the expected attention distribution under the null hypothesis of "no salient key." Every key receives equal weight. Nothing stands out.
Introduce one anomalous key. Let k1 have alignment e1 = E + δ with δ > 0, while the remaining N−1 keys stay at baseline E.
Mean logit: ē = [(E + δ) + (N−1)E] / N = E + δ/N
Logit residual for the salient key:
δ1 = e1 − ē = (E + δ) − (E + δ/N) = (N−1)/N · δ
The residual δ1 is proportional to δ, scaled by (N−1)/N. For large N, δ1 ≈ δ. The residual is the deviation from the uniform baseline.
Attention weight for the salient key:
α1 = exp(E + δ) / [exp(E + δ) + (N−1)exp(E)]
α1 = 1 / [1 + (N−1)exp(−δ)]
This is the exact closed form for the single-salient-key case. Three regimes:
Expand α1(δ) around δ = 0:
α1(0) = 1/N
α'1(0) = (N−1)exp(0) / [1 + (N−1)]² = (N−1)/N²
α1(δ) = 1/N + (N−1)/N² · δ + O(δ²)
Interpretation: For small deviations from uniform, the attention weight grows linearly with δ. The slope is (N−1)/N², which decreases with N — larger key sets require larger δ to achieve the same α1. This is the attention analog of the fact that a Class II anomaly (near-zero A, small ΔB) is harder to detect than a Class III anomaly (large negative A, large |ΔB|).
In the general case (non-uniform baseline, multiple keys with non-zero δj), the attention weight for key i is:
αi = softmax(δi) = exp(δi) / Σj exp(δj)
where δi = ei − ē is the residual.
Property 1': If all δj = 0, αi = 1/N for all i (uniform — "on the Mudrock line").
Property 2': If |δi| ≫ |δj≠i|, then αi → 1 (winner-take-most — "Class III bright spot").
Stage 1 — Extract: ej = q·kj/√dk (per-key alignment).
Stage 2 — Residual: δj = ej − ē, where ē = (1/N)Σj ej. Zero on uniform baseline, non-zero for salient keys.
Stage 3 — Amplify/Classify: αj = softmax(δj) = exp(δj)/Σi exp(δi). Exponential amplification of large residuals.
The amplification function is exponential softmax — stronger than AVO's linear threshold, same structural role. The signal is the residual, exponentially amplified.
| Stage | AVO | Attention | Structural Role |
|---|---|---|---|
| Feature Extraction | A (intercept), B (gradient) | ej = q·kj/√dk | Project raw observation into comparable coordinates |
| Baseline Model | Bbg(A) = mA + c | ē = (1/N) Σj ej | Expected value under null hypothesis (no anomaly) |
| Contrast Residual | ΔB = Bobs − (mAobs + c) | δj = ej − ē | Deviation from baseline — THE SIGNAL |
| Amplification | Θ(|ΔB| − τ) — threshold | softmax(δj) — exponential | Nonlinear function: zero near baseline, large for outliers |
| Output | Class I–IV anomaly flag | αj ∈ [0,1], Σαj = 1 | Anomaly classification / salience assignment |
Table 1: The complete stage-by-stage mapping. The contrast residual row (highlighted) is the core of the equivalence.
Both systems implement:
Output = Amplify( Observation − Baseline(Observation) )
where:
| AVO | Attention | |
| Observation | Bobs | ei |
| Baseline | Bbg(Aobs) = mAobs + c | ē = (1/N) Σj ej |
| Residual | ΔB = Bobs − Bbg(Aobs) | δi = ei − ē |
| Amplify | Θ(|ΔB| − τ) | exp(δi) / Σj exp(δj) |
The amplification stage differs in functional form — and this is where the analogy is strongest, not weakest:
The hard threshold's advantage is that it has a "dead zone" — small deviations produce zero output. Softmax has no dead zone. Every δi, no matter how small, contributes to αi. This is why ungoverned softmax inevitably hallucinates: it cannot produce a truly uniform output unless all inputs are exactly equal. Any floating-point variation, any noise, any slight embedding misalignment — softmax will amplify it into a non-uniform distribution.
This is the mathematical justification for Essay #12's central claim: physics-constrained attention is not optional; it is the only defense against softmax's inherent tendency to amplify noise.
The mapping degrades or breaks under these conditions:
Both domains exhibit a three-tier approximation hierarchy, trading exactness for tractability:
| Tier | Geophysics | Attention ML | Cost | Governance |
|---|---|---|---|---|
| Exact | Zoeppritz (1919) — 4×4 matrix, all mode conversions, complex beyond critical angle | Full softmax O(N²) — all key-query pairs, exponential, differentiable | Maximum | Physical constraint available (Zoeppritz) / can be added (external baseline) |
| Linearized | Aki–Richards (1980) — first-order in elastic contrasts, three additive terms | Linear attention (Katharopoulos et al., 2020) — φ(Q)φ(K)T, O(N), kernelized | Moderate | Partially governable; breaks under extreme conditions |
| Interpretable | Shuey (1985) two-term — A + B sin²θ, direct A–B crossplot interpretation | FlashAttention (Dao et al., 2022) — IO-optimized, tiled, exact output, approximate intermediates | Minimal | Governance MUST be external; physics not in the model |
Table 2: The parallel approximation hierarchy. As you descend, you gain speed and lose intrinsic physical fidelity. External governance must compensate.
The principle: the further you approximate, the stronger your external governance must be. Shuey two-term works brilliantly for Class III sands at moderate angles — and fails for Class IV. FlashAttention computes attention exactly but produces intermediates that cannot be audited. In both cases, the defense is an external baseline that the approximation cannot overwrite.
The mapping provides a formal language for describing what transformer-based seismic models are doing internally. An attention head that produces αi ≫ 1/N for tokens corresponding to a specific reflector is computing the equivalent of an AVO anomaly flag. The ACRisk framework (Essay #11) and the PCA framework (Essay #12) provide audit machinery.
The mapping reveals that softmax attention is an ungoverned contrast amplifier. It will amplify noise in the absence of signal. The AVO community's solution — an external, calibrated, falsifiable baseline — is directly portable. Every high-stakes transformer deployment needs a Mudrock line.
The contrast primitive — extract features, compute residual against baseline, amplify deviations — is a universal computational pattern. It appears in AVO, in attention, in the Pi-Transformer (Maleki & Pourmoazemi, 2025), in the Anomaly Transformer (Xu et al., 2022, ICLR), and in any system that must answer "what stands out from the background?" Recognizing this universality enables cross-domain transfer of governance tools, failure-mode taxonomies, and calibration methods.
TRILOGY VERDICT: The EUREKA announced in Essay #11, the governance demanded by Essay #12, and the derivation supplied by this paper (#13) together establish that contrast detection against a calibrated baseline is a universal computational primitive, instantiated in seismic AVO analysis and transformer self-attention with structurally identical three-stage computation. The difference in amplification function (linear threshold vs. exponential softmax) does not break the equivalence — it clarifies why ungoverned softmax hallucinates and why physics-constrained attention is the remedy. The governor is not optional. The Zoeppritz equations prove it can be built. The arifOS constitutional floors prove it generalizes. The trilogy is sealed.
By Arif Fazil Sealed 999 · 16 min read
Muhammad Arif bin Fazil
Geoscientist · Architect, arifOS · Petronas Carigali · UW–Madison '13
Penang, Malaysia
Published: 05 June 2026 · Direct Publication · /words/ context
Epistemic Tag: INT — interpretive synthesis across AI governance, institutional economics, and systems theory
Pairs with: Contrast-Governed Anomaly Detection (EUREKA I) · Physics-Constrained Attention: Zoeppritz (EUREKA II)