Live Portfolio — what to own today
| Date | Sleeve | Action | Ticker & name | Prior → New (% NAV) |
|---|
Performance
Cumulative return for the deployed blend and each of the four contributing sleeves, anchored at 0% at the start of the selected period. The stat strip and the attribution bars below the chart recompute live with the timeframe — switching to 1Y will show this year's return, switching to MAX will show full backtest history. Default YTD. Amber-shaded bands mark periods when the breadth regime overlay was active (RISK_OFF — half of NAV moved to SHY, 1-3y US Treasury).
blend_35_35_10_20_gated_eem_tilted
·
Universe A14 / B13 / C23 / D5
·
Weights 35 / 25 / 10 / 20 + EEM tilt 10%
·
Sharpe +1.31
·
CAGR +15.8%
·
Max DD −16.3%
Known caveats & methodology footnotes (acknowledged upfront)
- Survivorship bias in Strategy C (thematic): the 25 ETFs all survived to 2026. Failed thematics (cannabis, leveraged-thematic, volatility) never entered the universe. C standalone Sharpe is on a survivor-selected basket.
- Breadth regime gate parameters — partial OOS evidence: thresholds (off=20% / on=50% / derisk=50%) were chosen from a 12-variant in-sample sweep. June 2026 walk-forward validation (annual refit, same grid, 5y train + 1y test) confirmed the parameter surface is stable — refit picks similar thresholds (off=0.25, on=0.50) and matches the fixed deployed config almost exactly. However, on the 2024-2026 OOS test windows specifically the gate cost -0.13 Sharpe vs the un-gated blend, because those years didn't contain the major drawdowns (2020 COVID, 2022 inflation) the gate is designed to catch. The deployed +0.10 Sharpe contribution at the full-window level is therefore "insurance" economics — pays off rarely, bleeds slightly in between — not walk-forward-validated alpha. Treat the gate as drawdown protection conditional on future crashes resembling past ones, not as a Sharpe-lifting signal. See
data/risk_overlay_validation.jsonfor the per-segment numbers. - EM tilt — small-sample caveat partially closed by OOS test: the 50d/200d golden cross has 14 distinct ON-events in 7.5 years (~2/year). Earlier caveat was "deployed on weak sample"; June 2026 walk-forward validation showed the fixed deployed (50/200) thresholds contribute +0.06 Sharpe on the 2024-2026 OOS test windows vs the un-tilted blend, and annual refit picks slightly different windows (50/150) that perform marginally worse. The deployed thresholds are near-optimal — the OOS evidence is supportive, not conclusive, given the modest sample size. See
data/em_tilt_validation.json. - Sharpe sample noise: 7.5y × weekly observations gives Sharpe SE ≈ ±0.4. A reported Sharpe of +1.31 sits in a ~95% CI of roughly [+0.57, +2.05]. Treat headline ratios as ranges.
- No live track record: every return shown is simulated. Future performance will diverge.
- Backtest costs are conservative-not-pessimistic: 2–9 bps per unit weight change per sleeve, no slippage / market-impact model. Real execution at Friday close has measurable spread cost not captured here.
- Walk-forward definition: within-sleeve K parameter (the number of ETFs held) is refit each year using only past data, then applied to the next 12 months. Not walk-forward: weighting scheme, universe additions over time, per-sleeve cost calibration, breadth-gate parameters, EEM-tilt parameters. Those structural changes were applied across the entire backtest history — their reported lifts are in-sample improvements.
Combined holdings
Every position in the deployed portfolio, sorted by total weight. Each row shows the ETF, which strategy holds it (A = US sector breadth, B = asset-class momentum, C = thematic momentum, D = Europe sector breadth), why it's in the portfolio (the signal that put it there), the within-strategy weight, and the effective weight after the 35/35/10/20 blend. If you have $100,000 to deploy, the "Total weight" column tells you how much capital each ETF gets. Click any row to see the underlying ticker's 1Y price chart.
How "Total weight" is built, with one row as a worked example. Say IUES sits in Strategy A only with a within-strategy weight of 19.3%. Strategy A allocates 35% of total NAV, so IUES gets 0.35 × 19.3% = 6.76% of the combined portfolio. On $100k, that is $6,755 in XLE (the trade proxy for IUES). The four-strategy panels below show how each within-strategy weight is computed from the underlying breadth or momentum signal — hover any row for the exact arithmetic.
Exposure by asset class
Roll-up of the combined portfolio by broad category. Tells you what the strategy is actually betting on right now beyond individual ticker symbols.
Diversification check — realised return correlation between the four sleeves
Pearson correlation of daily returns (not breadth signals) between each unique pair of sleeves, computed on the common 2018-11 → today window. Lower = more diversifying — the blend's risk-adjusted return depends on these being well below 1.0. Lower-triangular layout shows only the six unique pairs (the diagonal of self-correlations and the redundant mirror entries are omitted).
Strategy A, B, C and D — side by side
The four engines that produced today's holdings. Strategy A picks the strongest US sectors by sector-relative breadth (each sector's breadth measured against the cross-sectional mean). Strategy B picks the strongest asset classes by ETF-level momentum across 12 broad ETFs. Strategy C is a small thematic sleeve for catching secular trends (AI, clean energy, biotech) across 25 themes. Strategy D is a Europe sector breadth sleeve for non-US macro orthogonality. Combined portfolio: 35% A + 35% B + 10% C + 20% D in normal markets; the EEM tilt reallocates 10pp from B to EEM when EEM/SPY relative-strength turns up (golden cross of the 50/200d ratio). Since Phase 29 (July 2026) this tilt is the portfolio’s only EEM position — EEM left B’s rotation universe after the WS2 review found the two roles double-counted.
How the deployment works in 6 steps
Recipe — combined 35/35/10/20 A+B+C+D, no leverage
Expected statistics (in-sample, common 2018-2026 window): Sharpe +1.15 (pre-overlay baseline), CAGR +15.1%, max drawdown -23.8%. Deployed Sharpe with the breadth regime gate and EM tilt active is +1.31 (note: the +0.16 Sharpe lift vs the +1.15 pre-overlay baseline is in-sample; OOS evidence: the EM tilt contributes +0.06 Sharpe in walk-forward, the regime gate is "insurance economics" — pays off on rare drawdowns like 2020 / 2022 but bleeds slightly in uptrending years; see the Caveats at the top of the page). If you prefer lower drawdown without the Europe sleeve, use the prior 45/45/10 A+B+C blend (Sharpe +1.09 pre-overlay, max DD -21.5%) — see Multi-Strategy tab.
The deployed 4-way blend's complete history — every position, every rebalance, every contribution. This tab consolidates the per-sleeve trade detail that lives on the four individual strategy tabs into one blend-level view. Useful for: walking a client through "how does this actually work over time", validating the performance attribution numbers, and answering "what did the portfolio hold during regime X" questions.
1 · Combined trade ledger — every rebalance across all four sleeves
Flat list of every position change across A, B, C, and D. One row per (sleeve × rebalance × ETF). Action = entry (first appearance) / exit (last appearance) / resize (weight change within an existing position). Filter by sleeve, ETF, or date to drill into a specific moment.
2 · Asset-class exposure over time
The deployed 35/35/10/20 A:B:C:D portfolio's exposure by asset class, stacked to 100%, weekly Friday snapshots since 2018-Q4 (common window across all four sleeves). Strategy A's 14 US sector ETFs are decomposed into 5 macro buckets (US Tech / Comms, US Cyclicals, US Energy, US Financials, US Defensives) so A's actual sector rotation is visible. Amber-shaded vertical bands mark periods when the breadth-regime overlay was RISK_OFF — look for the corresponding swell in the pale-green Cash (SHY) band at the bottom. Notice the bond spike in early-2020 (B's flight to TLT/IEF), the thematic burst in 2020-21 (C catching the AI/clean-energy boom), the commodity tilt in 2022, and the rotation between US Tech / Cyclicals / Defensives within Strategy A over the full window.
Bands sum to 100% of NAV every Friday. Colours grouped by exposure family: greens for cash + fixed income + real estate, blues / purples for US equities (defensives → cyclicals → tech), teals for international equities, amber for commodities, magenta for thematic optionality. Stack order bottom-up: defensive base first (cash, bonds, real estate), then US equities, then intl / commodities / thematic optionality — the "satellite" risk that moves most. Bucketing: Strategy A's sectors split into Tech / Comms (CNDX, SOXX, IUCM), Cyclicals (IUIS, IUCD, IUMS), Energy (IUES), Financials (IUFS), Defensives (IUHC, IUCS, IUUS); IUSP rolls into Real Estate; VGK aggregates with Strategy D's Stoxx 600 sectors into Europe Equity; Strategy C's commodity-equity themes (XME, GDX, COPX, WOOD, REMX, MOO) aggregate into Commodities; SHY (both sleeves' cash floor and the breadth gate's defensive parking) is its own band so the overlay engagement is visible.
3 · Blend-level attribution — total NAV contribution per ETF
Each ETF's total contribution to the deployed 35/35/10/20 blend's return, aggregated across the sleeves that held it. If an ETF appears in multiple sleeves (e.g., GLD in Strategy B AND Strategy C, SHY as cash floor in both B and C, JETS as thematic in C), its blend contribution is the sum of (sleeve_weight × sleeve_contribution) across each sleeve. This is the auditable per-ETF P&L attribution that proves the headline blend return is built from these specific position-level decisions. Sorted by absolute contribution.
Strategy A — US sector breadth rotation
Sector-relative breadth rotation, up to K = 7, long-only. Each Friday close, compute % above 200d MA for each of the 14 sector ETFs. Convert to sector-relative breadth by subtracting the cross-sectional mean. Rank by relative breadth, take the top K = 7, then drop any with negative relative breadth (those below the mean) to keep the sleeve long-only. Weight survivors by their share of total positive relative breadth. K is the cap, not a target — in narrow leadership regimes only a handful of sectors sit above the mean and the sleeve concentrates to those; in broad rallies all top 7 are above mean and the sleeve fans out. Refit K only annually on expanding-window Sharpe — the choice has been stable at K = 7 throughout the backtest. Strategy A is the largest single sleeve (35%) in the deployed multi-strategy blend; see the Multi-Strategy tab for how it combines with B, C, and D.
- Every Friday close: pull the 14 ETF rosters, compute % of constituents above their own 200d simple MA (absolute breadth, 0-100%).
- Convert to sector-relative breadth: for each ETF, subtract the cross-sectional mean breadth across all 14 sectors. The signal is now "how much is this sector leading or lagging the others" (positive = leading, negative = lagging).
- Rank by relative breadth and take the top K = 7, then drop any with negative relative breadth (sectors below the mean). The sleeve holds up to K survivors — often fewer when the cross-section is narrow.
- Weight by positive relative-breadth share: weight_i = rel_breadth_i / Σ(positive rel_breadths). Survivors with the largest excess over the mean get the largest slices.
- Rebalance to the new mix. The Trade Explorer deep-dive (below) shows the exact weights at every rebalance.
- Once a year, re-evaluate K ∈ {3, 5, 7}. So far the answer has been K = 7 every year.
Why fewer than 7? Phase 20.1 (deployed May 2026) switched the signal from absolute breadth ("% above 200d MA") to sector-relative breadth ("% above 200d MA, minus the cross-sectional mean"). The change empirically lifted Sharpe by +0.05 and cut SPY correlation from 0.93 to 0.83 — sectors moving with the market wash out, sectors actually leading on a breadth basis stand out. The long-only drop-negatives rule is what makes the sleeve self-thin in narrow regimes instead of holding 7 mediocre laggards.
Strategy A — current selection
What Strategy A actually holds right now, and why each sector qualified. The rule: rank every US sector ETF by sector-relative breadth (breadth − cross-sectional mean), keep only those above the mean, hold the top K (up to 7) weighted by positive excess share. The breadth shown here is the value at the last rebalance, not today — so a sector whose breadth has since crossed the mean is not added until the next rebalance, and a held sector that has since fallen below is not dropped until then either. That is why the Live Signal chart below may show a different ordering than this table.
Live signal — US Sector Breadth (Strategy A universe)
The live signal that drives every rebalance decision. Default: Leader vs Market — Strategy A's top-breadth holding plus CSP1 (S&P 500) as a broad-market reference. Click any chip to add or remove sectors from the chart. Use the timeframe selector above the chart to switch between 3M / YTD / 1Y / 3Y / 5Y / MAX windows, or drag across the plot to zoom into a custom date range (double-click to reset). Trend snapshot below the chart shows current breadth + change over recent past.
Trend snapshot
Current breadth + change over recent past. Positive Δ = strengthening; negative = weakening. Sorted by current breadth.
Portfolio construction verdict
Loading…
Headline comparison
Top row (highlighted green, tagged DEPLOYED) is the actual Strategy A used in the 4-way Multi-Strategy blend at 35% weight: Sector-relative breadth rotation, up to K=7, weekly Friday, no leverage. Below it: the K=5 unleveraged variant as a K-sensitivity reference, and the two passive benchmarks (naive equal-weight across the 14 ETFs, SPY buy-and-hold).
Equity curves — apples-to-apples
All five lines anchored at 0% at the common start date (the latest "data available" date across all series) so visual comparison is fair. Use the timeframe selector to view shorter windows.
Right-tail behaviour (extreme upside & downside metrics)
Metrics that Sharpe alone underrates. Sortino credits upside vol; skewness shows monthly-return asymmetry; rolling 12m extremes show the actual best/worst year experienced. Asymmetry ratio = |best| ÷ |worst| — values > 1 mean the upside tail dominates the downside tail.
Trade Explorer — Strategy A deep-dive
Every rebalance of Strategy A from — to —. Each row = a Friday close; holdings shown are what the strategy should be in for the following week. Green chips entered this week; red chips exited; blue chips persisted. Weight = portfolio share; breadth = % of that ETF's constituents above 200d MA at the rebalance date. For the BLEND-LEVEL combined ledger across all four sleeves, see the Trade History tab.
Strategy A · Equity vs benchmarks
Strategy A (K = 7, weekly Friday) versus SPY buy-and-hold and naive equal-weight across all 14 ETFs. All series shown as cumulative return from period start (0% baseline) — use the timeframe selector to switch between YTD / 1Y / MAX windows.
CAGR = compound annual growth rate. Δ = strategy minus benchmark.
Strategy A · Sector allocation over time
Weekly snapshot of the portfolio composition — stacked to 100%. Colours by ETF. Wide bands = persistent leadership; thin slivers = brief appearances. Tells you what the rotation was actually betting on in each market regime.
Strategy A · Trade history
All rebalances, newest first. Use the filter to find rebalances containing a specific ETF or close to a specific date.
Strategy A · Performance attribution
For each ETF, the contribution to Strategy A's total return. Contribution = sum of (yesterday's weight × today's ETF return) over all days in the backtest. % of total normalises across ETFs (positives net of negatives). Annualised return when held = the geometric daily return of just the ETF on days the strategy held it — answers "did the strategy time this sector well, or just own it through good periods?". Sorted by contribution.
Show all top-K variants (K ∈ {3,5,7} × weighting ∈ {equal, rank, breadth}, unleveraged)
Full unleveraged sweep — different ways to pick and weight the top-K-by-breadth basket. Weekly rebalance, 5 bps per unit of turnover.
Strategy B — Asset-class momentum
A separate strategy that operates at the asset-class level instead of within US sectors. 13 broad ETFs across US equity, international developed, emerging markets, real estate, commodities, and bonds. Each Friday close, rank by distance above own 200d MA. Hold the top K with positive signal, weight by signal strength; idle capital sits in SHY (1-3y Treasury) as a duration-neutral cash floor. IEF remains in the rotation universe as a long-duration play, not as the cash proxy. This is "which asset classes should I be in at all?" — orthogonal to Strategy A's "within US equities, which sector?".
- Each Friday close, compute (price − MA200) / MA200 for each of the 13 asset-class ETFs.
- Drop any ETF with negative signal (price below its 200d MA — in a downtrend).
- Rank survivors. Take the top K = 7.
- Weight by signal share. If only N < K ETFs have positive signal, allocate (N/K) of capital to them and the remaining (K − N)/K to SHY (1-3y Treasury, duration-neutral) as cash floor. Built-in cash floor when the world is broadly weak — unlike Strategy A which stays 100% invested.
- Rebalance to the new mix. Costs 5 bps × turnover.
Strategy B — current selection
What Strategy B actually holds right now, and why each name qualified. The rule: rank every asset-class ETF by how far it sits above its own 200-day average (the signal), keep only those above it, hold the top K weighted by signal share, and route any unfilled slots to cash. The signal shown here is the value at the last rebalance, not today — so a name whose live signal (see the Trend snapshot below) has since overtaken a held one is not added until the next rebalance. That is why an ETF can show a higher current signal yet still read "no" in that snapshot.
Live signal — Asset class momentum (Strategy B universe)
The live momentum signal that drives Strategy B rebalance decisions, across the 13 asset-class ETFs. Y-axis is distance above each ETF's own 200d MA (can be negative); the dashed guide at 0% is B's eligibility line — only names above their own 200d MA are ranked. Default: today's Strategy B leader + SPY as reference. Use the chips to add or remove ETFs, and the timeframe selector above the chart to switch between 3M / YTD / 1Y / 3Y / 5Y / MAX windows (drag across the plot to zoom into a custom date range, double-click to reset). The y-axis rescales to the selected window. Strategy C applies the same signal definition to its own 25-ETF thematic universe under a stricter floor — that chart now lives on the Strategy C tab.
Trend snapshot
Current momentum signal + change over recent past. Sorted by current signal value.
Equity vs benchmarks (18-year backtest)
Strategy B versus SPY buy-and-hold (broad-market passive — note its catastrophic -51% drawdown in 2008), the 60/40 SPY/IEF balanced portfolio (the conventional benchmark), and naive equal-weight across all 14 ETFs (diversification only, no signal). Strategy B's longer history (back to 2008) catches the GFC, the 2011 mini-crash, 2015-2016 commodity bust, 2018 Q4, COVID 2020, 2022 rates shock, and the recoveries in between. The headline result is the massive drawdown reduction.
Right-tail behaviour (extreme upside & downside metrics)
B is the boring sleeve and that is the point — narrowest distribution, lowest 12m extremes, smallest skewness, top sleeve only ~10% of months. The asset-class rotation does not catch fads; it rotates into bonds when equities weaken. Its asymmetric value is downside protection, not upside capture.
Asset-class allocation over time
Weekly stacked-to-100% allocation. Notice the regime shifts — the strategy rotated decisively into bonds (TLT/IEF) in 2008 and 2020, into commodities during the 2022 inflation shock, and into equities through the 2017-2019 and 2023-2024 expansions. This regime-following behaviour is what produces the drawdown reduction.
K × cadence sensitivity
Same heat-grid format as Strategy A's Test 11. Best Sharpe cell highlighted green, worst red.
Walk-forward K refit
Annual refit picking K ∈ {3, 4, 5, 6, 7} on expanding-window Sharpe, applying that K to the next 12 months. Same methodology as Strategy A's Test 10.
Trade history
Newest first. Filter to find specific ETFs or dates.
Performance attribution
Per-ETF contribution to the strategy's total NAV. The signal column shows the asset class — tells you which classes drove returns across the 18-year window.
Strategy C — Thematic sleeve
A third sleeve for catching secular trends that don't fit traditional sectors (AI, cybersecurity, clean energy, biotech, blockchain, defence, crypto, broad metals & mining, timber, rare earth, China-tech-broad, China-A-share semis, etc). Same momentum signal as Strategy B applied to 25 thematic ETFs, with extra guardrails to limit fad-chasing. Sized at 10% of the combined portfolio — designed as optionality on the next AI-style bull run, not a primary alpha source.
- Each Friday close, compute distance above 200d MA for each of the 25 thematic ETFs.
- Hard signal floor: only consider ETFs with signal ≥ +5% above their 200d MA (not just positive). This filters out marginal "in an uptrend" cases that often reverse.
- Sleeve breadth gate (deployed June 2026): if fewer than 30% of the 25-theme universe clears the +5% floor at this rebal, exit ALL positions to SHY for the week. Catches sleeve-wide regime changes (e.g. early 2022 thematic-complex rollover) before per-ETF signals individually breach. Walk-forward Sharpe lift +0.16, max DD reduction 14.8pp vs the un-gated rule.
- Rank the survivors. Take the top K = 5. Walk-forward picked K=5 in every refit segment once the sleeve gate was active — the gate handles regime risk so the in-sample concentration choice K=5 becomes safe to deploy. (Prior K=4 was the WF pick before the gate.)
- Equal-weight (1/K) across the top K. The +5% signal floor already filters out modest trends, so signal magnitude beyond eligibility carries little extra information — every eligible candidate is well into an uptrend. Equal-weighting prevents the most-overbought ETF (statistically the one most likely to mean-revert) from being overweighted.
- Cash floor: when fewer than K candidates clear the +5% threshold, the deficit sits in SHY (1-3y Treasury, duration-neutral cash equivalent) as cash proxy.
Strategy C — current selection
What Strategy C actually holds right now, and why each theme qualified. The rule: rank every thematic ETF by distance above its own 200-day MA, keep only those at or above the +5% hard signal floor, then equal-weight (1/K) across the top K=5. If fewer than K names clear the floor, the deficit sits in SHY. The V6 sleeve-breadth gate trips when fewer than 30% of the 25-theme universe is above +5% — when active, the entire sleeve exits to SHY for the week. The signal shown here is the value at the last rebalance, not today — a name whose live signal has since crossed +5% or fallen below it is not acted on until the next rebalance.
Live signal — Thematic momentum (Strategy C universe)
The live momentum signal that drives Strategy C rebalance decisions, across the 25 thematic ETFs plus SHY (the cash proxy). Same signal definition as Strategy B — distance above each ETF's own 200d MA — applied to a different universe under a stricter rule, so two guides are drawn: 0% marks the 200d MA crossing, and the dotted line at +5% is C's hard signal floor. That makes the eligible set readable at a glance, and with it the distance to the 30%-of-universe sleeve-breadth gate — use the "Above +5% floor" chip action to select exactly the themes clearing it today. Default: today's Strategy C holdings + SPY as reference (SPY is not a thematic candidate; it is drawn from Strategy B's series as the sleeve's benchmark, clipped to C's own history). Drag across the plot to zoom into a custom date range, double-click to reset. Series begin 2018-11 with the sleeve's backtest, except 159801.SZ which lists from 2020-11.
Trend snapshot
Current momentum signal + change over recent past. Sorted by current signal value. The current signal column marks green at or above C's +5% floor, not at 0% — eligibility here needs the floor, not merely an uptrend.
Honest assessment — Strategy C is an optionality sleeve, not a Sharpe-alpha sleeve
The standard quant lens (Sharpe ratio, bootstrap p(better) — the estimated probability the strategy beats the alternative on a future 7.5-year sample) systematically underrates Strategy C. Sharpe penalises upside vol symmetrically with downside vol; bootstrap p(better) measures the mean outcome, not the tail outcome. C is structured as an optionality sleeve — capped 10% sleeve weight (limits downside per year), unbounded upside if a thematic bull fires. The right metrics are right-tail metrics.
The empirical case for C, on right-tail metrics (see "Right-tail behaviour" section below):
- Best rolling 12-month return: +162%. Vs A +85% / B +43% / D +66%. C delivers the largest absolute upside tail in the universe by 2-4×.
- Top-performing sleeve 41% of months (most of any sleeve). When C wins it wins often and big; the bootstrap-on-Sharpe missed this because magnitudes are highly variable.
- COVID + thematic boom (Mar 2020 → Feb 2021 ARKK peak): C delivered +170% standalone. The 50/50 A:B got +58% in the same window. The 4-way blend with 10% C got +65% — the small sleeve captured ~70% of C's idiosyncratic alpha.
- 2022 inflation crash: C was -5.4% (with V6 sleeve-breadth gate active, which moved C to SHY for most of 2022; without the gate C would have lost -29%). The gate prevented what would otherwise have been the C sleeve's worst year contribution. Blend damage in 2022 was -9% (vs SPY -19%).
The asymmetry is the point. At a 10% sleeve weight, C contributes ~+17pp to blend return in its best year and ~-2.5pp in its worst year — a 7:1 upside/downside ratio. That is textbook optionality. Backtests of the last 7.5 years include the 2020-21 boom but cannot price the OPTIONS embedded in C — its real value compounds whenever a new thematic regime emerges that we cannot predict today.
Caveat: walk-forward Sharpe is +0.61 vs in-sample +0.79 (canonical 25-ETF universe at K=5 with the sleeve-breadth gate deployed June 2026; was +0.45 / +0.80 without the gate at K=4 — the gate lifts WF Sharpe by +0.16 by exiting the sleeve to SHY during sleeve-wide regime changes like the 2021-22 thematic blow-up. The 14.8pp max-DD reduction comes with a documented V-shape downside: the gate locks in the bottom during fast COVID-style recoveries — see data/thematic_exit_robustness.json for the full episode-by-episode attribution) — modest degradation, which is structural to thematic momentum. The 10% sleeve cap is the risk-management response to this. Do not deploy C as a primary alpha source; do deploy it as a small optionality sleeve.
Equity vs SPY (7.5-year backtest, 2018-11 → today)
Strategy C versus SPY buy-and-hold as the obvious passive benchmark. Note the deep drawdowns (-48% peak-to-trough after the Bitcoin addition; was -43% pre-BTC) — thematic ETFs cluster on tech, clean-energy, and now spot crypto, all of which run high vol. The recovery has been strong but at high vol. The deep drawdowns are the option premium; the +170% COVID-era return and 2024 crypto rally are the option payoffs.
Right-tail behaviour — the optionality scorecard
This is the section that earns Strategy C its place in the deployed blend. The asymmetry ratio (best 12m ÷ |worst 12m|) and the % months as top sleeve are the metrics that capture optionality value. Compare against the other strategy tabs to see how C's right tail dwarfs everything else.
Thematic allocation over time
Weekly stacked-to-100% allocation. When the cash floor activates (fewer than K = 5 themes clear the +5% threshold) the deficit sits in SHY. When the sleeve-breadth gate fires (fewer than 30% of universe above +5%) ALL allocation goes to SHY for the week — these are the big monolithic SHY bands during regime changes. Notice the rotation through commodity-equity (gold miners, uranium, copper, lithium) when tech was weak in 2022.
K × cadence sensitivity
Same heat-grid format as Strategy A's Test 11 and Strategy B's grid. Best Sharpe cell highlighted green, worst red.
Walk-forward K refit
Annual refit picking K ∈ {3, 4, 5} on expanding-window Sharpe. The walk-forward Sharpe is significantly lower than the in-sample Sharpe — a known characteristic of thematic ETFs (recent winners get extrapolated; the rotation chases fads). This is one reason for the 10% sleeve cap.
Trade history
Newest first. Filter to find specific themes or dates.
Performance attribution
Per-ETF contribution to Strategy C's total NAV. The theme column shows which thematic categories drove returns.
Strategy D — Europe sleeve
A fourth sleeve applying the same constituent-breadth mechanism as Strategy A, but to 5 Stoxx Europe 600 sector UCITS funds (Banks, Oil & Gas, Technology, Industrials, Utilities). The motivation is structural orthogonality: Europe runs on its own macro cycle (ECB rates, EUR/USD, China trade exposure) that is genuinely different from the US sector universe. Sized at 20% of the combined portfolio in the recommended 4-way blend, which empirically lifts Sharpe by ~+0.07 vs the prior 3-way 45/45/10 baseline.
- Each Friday close, compute % of constituents above 200d MA for each of the 5 Europe sector ETFs.
- Rank by breadth. Take the top K = 3.
- Weight by breadth-share excess (each holding's weight ∝ its breadth − the K+1-ranked ETF's breadth). Same weighting rule as Strategy A.
- Rebalance to the new mix. Costs 5 bps × turnover.
- Trade as: the underlying Xetra-listed UCITS (EXV1.DE etc.) in EUR. Settlement T+2, full liquidity through any EU broker.
Strategy D — current selection
What Strategy D actually holds right now, and why each Europe sector qualified. The rule: rank the 5 Stoxx Europe 600 sector UCITS by absolute breadth (% of constituents above their own 200-day MA), hold the top K=3 weighted by breadth share. The 5-sector universe is small enough that all three slots are always filled — there is no cash floor. Unlike Strategy A, D uses absolute breadth rather than sector-relative excess because the cross-sectional mean over only 5 sectors does not add reliable signal. The breadth shown here is the value at the last rebalance, not today — a sector whose breadth has since changed is not re-ranked until the next rebalance.
Live signal — Europe Sector Breadth (Strategy D universe)
The live signal that drives every Europe rebalance decision. Default: today's Strategy D leader + the universe-wide mean as reference. Same metric as Strategy A (% of an ETF's constituents above their own 200d MA, 0-100%) but for the 5 Stoxx Europe 600 sector UCITS. Click any chip to add or remove sectors from the chart. Use the timeframe selector above the chart to switch between 3M / YTD / 1Y / 3Y / 5Y / MAX windows, or drag across the plot to zoom into a custom date range (double-click to reset). Note this series is daily, whereas Strategy A's breadth panel is weekly Friday — the same window holds ~5× the observations here.
Trend snapshot
Current Europe sector breadth + change over recent past. Positive Δ = strengthening; negative = weakening. Sorted by current breadth.
Honest assessment — Strategy D is a Sharpe lifter but raises drawdown
Standalone Sharpe over the common 2018-11 → 2026-05 window is +0.93 — better than SPY's +0.77 and similar to Strategy B's +0.95. CAGR is +14.9% with max DD -32.0%. The DD is fully participatory in the 2020 and 2022 sell-offs; Europe sectors did not provide downside protection.
When blended at 20% with the existing 45/45/10 A:B:C baseline (35/35/10/20 A:B:C:D), Sharpe rises from +1.08 to +1.15 and CAGR from +14.9% to +15.1%, but max DD widens from -21.5% to -23.8%. This is not a free lunch — it is a real diversification trade: marginally better risk-adjusted returns at the cost of ~2.3pp deeper drawdowns. If drawdown floor is your priority, stay with the 3-way 45/45/10. If Sharpe is your priority, switch to 35/35/10/20. (Sharpe figures here are pre-overlay baselines on the common 2018-2026 window. The currently deployed blend with the breadth regime gate and EEM tilt active has Sharpe +1.30.)
Pipeline fix note: the `compute_ma200_breadth` function originally used `rolling(200, min_periods=200)` which dropped the MA for any constituent with even 1-2% missing days. US constituents (S&P 500 via yfinance) have ~100% coverage so this was a no-op for Strategy A. Non-US constituents (.L / .DE / .PA / .AS / .MI) have sparse missing days from local holidays / dividend events, which caused breadth to silently freeze on a stale value (last good was 2023-04-06 in the broken version). Fix: relaxed `min_periods` to 90% of the window, allowing the typical sparse missingness. Strategy A Sharpe changed by <0.01 (US is the universe where the fix is a no-op). All numbers above are on corrected data.
Equity vs benchmarks (8-year backtest, 2018-01 → today)
Strategy D versus VGK buy-and-hold (Vanguard FTSE Europe ETF — USD-denominated, ~85% of European market cap, the like-for-like Europe-broad comparator) as the primary benchmark, with SPY retained as a cross-region US reference. The early flat period before 2018-Q2 reflects the 200d MA warmup; from there forward Europe rotation tracks above VGK on a Sharpe basis. The delta columns below are versus VGK.
Right-tail behaviour (extreme upside & downside metrics)
Strategy D's profile is closer to A than to C — solid rolling 12m upside, modest skewness, top sleeve about 28% of months. The optionality comes from regime orthogonality vs US sectors (ECB / EUR / China cycle), not from thematic convexity.
Europe sector allocation over time
Weekly stacked-to-100% allocation across the 5 Stoxx Europe 600 sector ETFs. The rotation visibly shifts between defensives (Utilities) and cyclicals (Banks, Industrials) as the European macro regime changes.
K × cadence sensitivity
Same heat-grid format as Strategies A / B / C. K = number of Europe sectors held; cadence = how often we rebalance. Best Sharpe cell highlighted green, worst red.
Trade history
Newest first. Filter to find specific sectors or dates.
Performance attribution
Per-ETF contribution to Strategy D's total NAV. The sector column shows which Stoxx 600 sectors drove returns.
Risk Overlay — defensive layer on top of the 4-way blend
A breadth-based regime gate that scales the deployed portfolio defensively when the broad US market is structurally weak. It is the difference between the un-gated 4-way blend and the deployed gated portfolio you actually hold. The gate fires on a single observable: the share of S&P 500 constituents above their own 200d moving average. When this breadth falls below the regime threshold the blend de-risks to a 50/50 mix of the unchanged active blend + SHY (1-3y Treasury); when it crosses back above, it returns to fully-invested. The thresholds (off=20% / on=50% / derisk=50%) were chosen in-sample but walk-forward validated June 2026 — the parameter surface is stable.
Honest framing — insurance economics, not alpha. The gate's full-sample contribution of +0.10 Sharpe and +7.5pp max-DD reduction comes overwhelmingly from two events: the 2020-Q1 COVID crash and the 2022 inflation bear. On the 2024-2026 walk-forward OOS test windows the gate cost -0.13 Sharpe vs the un-gated blend because those years contained no major drawdowns. This is the structural shape of a rare-event insurance policy — pays off massively during crashes, bleeds slightly between them. Deploy it for drawdown protection conditional on future crashes resembling past ones, not as a walk-forward-validated alpha signal. Per-segment numbers in data/risk_overlay_validation.json.
EM Tilt — emerging-markets cycle-turn overlay
A second mechanical overlay that operates independently of the breadth regime gate. Tilts 10% of the blend into EEM (emerging markets) when EEM/SPY relative strength turns up — specifically when the 50d/200d ratio crosses up through 1.0 (golden cross of the relative-strength line). Funded from Strategy B (35% → 25% during tilt-ON). The position is structural: the strategy holds EEM continuously while the cross remains positive, and unwinds back to baseline weights when the relative-strength line crosses back down. Since Phase 29 (July 2026) this tilt is the portfolio's ONLY EEM expression — EEM is no longer a Strategy B rotation member; the WS2 review found the two roles double-counted (look-through EEM peaked at 15% of NAV with both roles holding it simultaneously on 26% of days).
OOS-supported. Unlike the Risk Overlay's rare-event insurance economics, the EM Tilt fires often enough (~14 distinct ON-events in 7.5 years, ~2/year) that walk-forward test windows do contain the events the signal is designed to catch. June 2026 walk-forward validation found the fixed deployed (50/200) thresholds contribute +0.06 Sharpe vs the un-tilted blend on the 2024-2026 OOS test windows. Annual refit picks slightly different windows (50/150) but performs marginally worse than fixed — the deployed thresholds are near-optimal on the parameter surface. Per-segment numbers in data/em_tilt_validation.json.
Loading EEM tilt data…
Multi-strategy combination — A, B, C, D
Strategy A (US sector top-K-by-breadth, Sharpe ~0.98) handles within-US-equity rotation. Strategy B (asset-class top-K-by-momentum, Sharpe ~0.95 with max DD only -14%) controls drawdown via flight-to-bonds in crises. Strategy C (thematic top-K-by-momentum) adds optional exposure to secular trends (AI, clean energy, etc). Strategy D (Europe sector top-K-by-breadth, Sharpe ~0.93) adds structural orthogonality from a non-US macro cycle. Blended fixed-weight, rebalanced weekly. The deployed default is 35/35/10/20 A:B:C:D — the 4-way winner. See the verdict below before choosing.
Where the deployed performance comes from — full attribution
Each component's contribution to the deployed blend across three dimensions: CAGR (how much return it adds), Max drawdown (what it brings to the standalone DD profile, or for overlays what it removes), and Sharpe (its risk-adjusted footprint). This answers why we have the two overlays: Risk Overlay sacrifices a small slice of CAGR to take material chunks off the worst drawdown; EM Tilt's contribution is intentionally small (it is a low-cost positional bet on a cycle turn, not an alpha source).
Contribution by sleeve over time — and vs benchmarks
Each pastel-shaded layer is a sleeve's weighted contribution to the deployed blend's cumulative return at time t: wi × (equityi(t) / equityi(0) − 1). The thick red line is the deployed blend (with Risk Overlay + EM Tilt) — the dominant reference. The dashed near-black line is the ungated blend (sleeves only, no overlays); the gap between red and near-black shows the overlay impact directly. Two passive benchmarks for context: dotted brown = SPY buy-and-hold, dotted dark teal = 60/40 SPY/IEF.
Compare other blend variants — full side-by-side equity curves
Comparison view across every blend variant tested: 4-way A:B:C:D, 3-way A:B:C (pre-Europe baseline), 2-way A:B, the meta-rotation, plus each sleeve standalone. Default view shows only four lines — the deployed 4-way blend, the prior 3-way baseline, the 2-way 50/50 A:B (for "does adding C help?"), and Strategy A standalone. Click any legend chip below to add or remove a line; the chart redraws immediately.
Regime decomposition — how does each blend behave in specific market environments?
Per-strategy and per-blend total return + max drawdown across four hand-picked regimes from the backtest window. This is where Strategy C earns its place in the deployed blend — the COVID + thematic boom row tells the optionality story most directly. Each cell shows total return; sub-cell shows max DD.
Statistical significance — is the deployed blend distinguishable from the alternatives?
The Sharpe improvements documented across the design evolution are point estimates. To know whether they are real signal or sample noise, we paired-bootstrap the daily return series (moving block bootstrap, block 60d, 2,000 samples). Each row below shows the Sharpe differential point estimate, the 95% bootstrap CI on the differential, and p(better) — the fraction of bootstrap samples where the deployed blend's Sharpe exceeded the alternative. If the CI's lower bound is positive, the improvement is statistically significant at the 5% level (one-sided).
Caveat — what Sharpe bootstrap misses: Sharpe ratio treats positive and negative volatility symmetrically. For optionality sleeves like Strategy C, this systematically underrates the strategy's value. The "C does not lift the blend" finding above is true on average but misses C's asymmetric upside contribution in specific regimes. See the Regime decomposition section above for the actual right-tail evidence: in the COVID + thematic boom window (Mar 2020 → Feb 2021), Strategy C standalone delivered +170%; the 4-way blend's 10% C sleeve added ~7pp to blend return in that window vs no-C blends. The 10% sleeve cap is the optionality structure — bounded downside, unbounded upside — which is precisely the shape of bet that Sharpe-based metrics cannot price properly. The C decision is justified on right-tail / regime metrics, not on bootstrap Sharpe.
How to choose your blend — and what the data actually says
Sharpe figures below are pre-overlay baselines from the original blend selection (on the common 2018-2026 window). The currently deployed blend with the breadth regime gate and EEM tilt active has Sharpe +1.30 — see the configuration ribbon at the top of this page.
If you want best risk-adjusted return (default): 35/35/10/20 A:B:C:D. Highest Sharpe of any variant tested (~+1.15 pre-overlay on the common 2018-2026 window). The Europe sleeve adds genuine diversification from a non-US macro cycle (ECB, EUR/USD, China trade). Cost: max DD widens by ~2.3pp vs the 3-way baseline because Europe is real equity exposure with real volatility — not a downside hedge.
If you want lowest drawdown with thematic optionality: 45/45/10 A:B:C (the prior 3-way baseline). Sharpe ~+1.09 pre-overlay (slightly worse than 4-way) but max DD ~-21.5% (~2.3pp shallower than 4-way). The right choice if drawdown floor matters more than incremental Sharpe.
If you want lowest drawdown overall: 30/70 A:B (no C, no D). Sharpe still strong (~+1.10 pre-overlay), CAGR ~+13%, max DD only ~-17.7%. Cleanest two-sleeve construction.
If you want maximum CAGR and tolerate big drawdowns: Strategy A alone (~+17.5% CAGR but -31% max DD). The blends trade ~2-3pp of CAGR for ~7-10pp of drawdown reduction — usually a good trade.
The meta-rotation variant (own only A or only B each week, picking by trailing 6-month Sharpe) performs worse than either strategy alone. The lookback lags too much; by the time it identifies that B is winning, A has often started winning again. Static blends beat dynamic blends in this dataset.
How to think about Strategy C: C is an optionality sleeve, not a Sharpe-alpha sleeve. Bootstrap-on-Sharpe gave C ~42% p(better) vs no-C blends — true on AVERAGE outcome, but misses the asymmetric option payoff. C's right-tail metrics tell the real story: best rolling 12m +187% (vs A +85% / B +43% / D +66%), top performer in many months when thematic regimes turn, and in the COVID + thematic boom regime delivered +173% standalone (capturing ~70% of which got into the 10% blend sleeve). The 10% sleeve cap structures it as a long-dated out-of-the-money call basket — small premium (the -5.4% it lost in 2022 with V6 active was -0.5pp at the blend level; without V6 the loss would have been ~-2.9pp), unbounded upside (next AI/clean-energy/space/quantum boom). See the Regime decomposition section above for the empirical evidence and the Strategy C tab for the full optionality scorecard.
Caveat on Strategy D: the Europe sleeve is a Sharpe lifter but it raises max drawdown — Europe sectors are participatory in global equity sell-offs (2020, 2022), not defensive. The 4-way trade is "marginally better risk-adjusted return for ~2pp deeper drawdowns". If you primarily care about downside protection, stay with 45/45/10. If you care about Sharpe and CAGR, switch to 35/35/10/20.
Risk & Validation — why we deploy what we deploy
Loading…
Why top-K rotation, not per-ETF threshold tuning
The same MA200 breadth signal can be deployed three different ways. Each paradigm has a different number of free parameters, a different overfit profile, and a different walk-forward Sharpe. This is the strategic justification for the current Strategy A architecture: cross-sectional top-K rotation over per-ETF L-threshold tuning.
Walk-forward K selection — per-segment detail
For the deployed cross-sectional rotation paradigm, K (number of top-breadth ETFs to own) is refit annually on the expanding train window. Each refit picks the K ∈ {3, 5, 7} that maximised Sharpe so far, then applies it to the next 12 months. Concatenate to get the realistic OOS curve. The K sequence has been stable, which is itself a robustness signal.
Strategy A — K × rebalance-cadence sensitivity
How sensitive is the deployed Strategy A to its two free choices: how many ETFs to own (K) and how often to rebalance? Each cell is the in-sample Sharpe; sub-cell shows max drawdown, annual turnover (a unit = full portfolio replacement), and number of position flips. Best cell highlighted green, worst red. The deployed cell is K = 7 × Weekly Fri.
Universe correlation diagnostic — is the universe saturated?
Pairwise Pearson correlations of the underlying signal time series across each strategy's universe. Diagnostic for the question "would adding more ETFs help?" — if existing ETFs already cluster at correlations > 0.85, the universe is saturated and the marginal added ETF just adds turnover without new signal. If correlations are mostly < 0.5, there is room to diversify the universe.
Strategy A — breadth signal correlations
14 US sector/broad ETFs (after pruning IUIT in May 2026). Each cell = Pearson correlation of weekly breadth series. Red = highly correlated (redundant). Blue = uncorrelated (diversifying).
Strategy B + C — momentum signal correlations
13 asset-class + 25 thematic ETFs (deduped + SHY cash floor). Same metric — Pearson on weekly distance-above-200d-MA series.
The strategy explained, plain English
Own the US sectors that are leading the others on breadth — up to seven of them — each week. Skip the laggards. Update weekly. That is the whole strategy.
What is a "sector"? The US stock market is split into industry groups — Financials, Energy, Health Care, Industrials, Consumer Discretionary, Consumer Staples, Utilities, Materials, Communication Services, Real Estate — plus a few broad-market and concentrated picks (S&P 500, NASDAQ-100, Semiconductors, Small-cap 600). We use 14 ETFs that represent these groups. (NASDAQ-100 already gives us the large-cap tech exposure; we pruned the S&P 500 Info Tech ETF in May 2026 because it duplicated the NASDAQ-100 too closely.)
How do we judge which sector is "strong"? Every sector contains 30 to 500 individual companies. We count what percentage of those companies are currently trading above their own 200-day moving average. A stock above its 200-day average is in an uptrend; below it, a downtrend. So if 85% of the companies in a sector are in uptrends, that sector is strong — broadly strong, not just lifted by a few mega-caps. We call this number the breadth of the sector. It ranges from 0% (everything broken) to 100% (everything in an uptrend).
What is "leading on breadth"? A sector at 75% breadth sounds strong in absolute terms, but if every other sector is also above 75% then this one is not standing out — it is just riding the broader rally. To distinguish leaders from market beneficiaries, we compute the average breadth across all 14 sectors, then subtract that average from each sector's own breadth. The result is sector-relative breadth — positive when the sector is leading the pack, negative when it is lagging. This is the signal we actually rank on.
What this strategy is NOT doing. It is not trying to time the market — there is no "stay in cash if everything looks bad" rule. It is always 100% invested. It is also not trying to predict which individual sector will go up next. It is doing something simpler: own whatever is leading, trim whatever is lagging. Markets reward this pattern more often than they punish it, because trends in market breadth tend to persist for weeks-to-months at a time.
Why we trust it. The "Robustness" tab shows the strategy holds up when we apply the standard quant-research stress tests: walk-forward validation (refitting the only free parameter, K = the number of sectors to own, each year on out-of-sample data), bootstrap confidence intervals, sub-period decomposition through bear markets, sensitivity to the moving-average lookback and the rebalance frequency. The walk-forward Sharpe ratio is approximately 1.05 with no degradation from the in-sample number — that is unusual and is the main reason we picked this paradigm over per-ETF threshold tuning or fixed-threshold timing.
Caveats and what would kill it: a regime where all 14 sectors fall together (2008-style systemic crisis) hurts the strategy because owning the "strongest" of a falling universe is still owning falling stuff. We do not currently hedge or shift to cash. A modest improvement on the to-do list is to add a "if median breadth across all sectors is below 40%, reduce gross exposure to 50%" overlay — but we have not yet validated it OOS so it is not in the headline numbers.
Investment universe & selection — visual summary
For each sleeve: what could be held (universe), what filter decides which subset, what is actually held (selection). Universe counts and current-selection counts auto-populate from the live JSONs. Two overlays sit on top of the 4-sleeve blend — the breadth-regime gate and the EM tilt — each has its own dedicated tab (Risk Overlay · EM Tilt).
Why this universe? — design rationale per sleeve
Strategy A · US sectors
The 14-ETF mix decomposes into 9 GICS sectors (the iShares UK S&P 500 sector UCITS slate), 1 REIT proxy (IUSP — no iShares UK Real Estate sector UCITS exists, so substituted), 3 broad indices (CSP1, CNDX, IDP6) which act as the “no specific sector is leading” fallback in undifferentiated breadth regimes, and 1 sub-industry (SOXX — semiconductors). SOXX is the only non-sector inclusion: a discretionary prior-belief pick based on the long-standing practitioner observation that semis lead the tech cycle by 1–3 months (Philadelphia SOX has been followed as a cyclical bellwether since the 1990s). IUIT (S&P 500 Info Tech) was pruned in May 2026 because of 0.97 correlation with CNDX — no incremental information, just turnover noise. The Phase 5 retrospective tested 11 sub-industry ETFs (XME, IGV, XRT, FDN, XOP, OIH, KIE, ITB, AMLP, PHO, KRE) as thematic-sleeve additions, not direct Strategy A candidates — 4 failed the within-C correlation gate, 3 failed the cross-strategy gate (cousins of A's existing sectors), and 4 survivors added together degraded Strategy C's walk-forward Sharpe by 0.10. Reverted. See the Phase 5 research-log accordion below for the full per-ETF correlation table. Hindsight-bias check on SOXX (2026-05-30): we retroactively ran SOXX through the symmetric “all eligible sub-industries” gate it never originally got (see scripts/run_strategy_a_universe_gate.py). Result: SOXX passes all four criteria — max correlation with the other 13 ETFs is 0.83 (vs CNDX, just below the 0.85 threshold); in-sample K=7 Sharpe lifts +0.06 with SOXX added; max drawdown is essentially unchanged; and the walk-forward K refit picks K=7 every year with SOXX in the universe vs K=3 every year without it (the cross-section is too correlated to deploy at K=7 without SOXX's orthogonal information). The discretionary prior-belief inclusion turned out empirically defensible. IBB (biotech) is queued for the same gate test — pending availability of constituent breadth data via either iShares Europe biotech UCITS or NASDAQ Biotech Index constituents from a non-blocked source (iShares US holdings endpoint is currently Akamai-blocked). SOXX operational note (2026-05-31): the same iShares US block affects SOXX itself — SOXX is the only Strategy A member sourced from the US endpoint (all others use iShares UK). When the block engages, fetch_constituents.py falls back to carrying the most recent known-good roster forward (the PHLX Semiconductor Index changes ~2 holdings per year, so a 2-4 week stale roster has negligible signal impact). Breadth values are still computed daily against fresh prices — only the roster snapshot is stale. The weekly CI workflow re-runs this fetch + carry-forward + breadth recompute, so SOXX is never silently dropped from the eligible universe again (the issue spotted on 2026-05-31 was that the constituent fetch had not been run for 17 days, so the staleness guard masked SOXX out of three consecutive rebals). Data integrity policy (2026-05-31): a hard staleness ceiling — if any constituent roster crosses 30 calendar days since its last real fetch, fetch_constituents.py exits code 2 and pipeline.py aborts the dashboard publish entirely (the staleness policy, escalation procedure, and incident log live in DATA_INTEGRITY_POLICY.md at the repo root). A non-fresh roster (15-30 days stale) shows as a yellow banner under the Live Portfolio hero card on this page. Secondary roster source (2026-05-31): SEC EDGAR N-PORT-P as a registered secondary source for SOXX (scripts/edgar_nport.py, ~400 lines + 10 tests). N-PORT-P is the quarterly regulatory filing that every US-registered ETF (including SOXX) submits to the SEC under Rule 30b1-9 — same data, more authoritative source. When the iShares US primary fails AND the EDGAR snapshot is fresher than the carry-forward source, the fetcher injects the EDGAR roster automatically. CUSIP-to-ticker resolution via OpenFIGI free tier (10 mappings per request, on-disk cache). Quarterly cadence means worst-case staleness from EDGAR alone is ~150 days; combined with daily breadth recompute against fresh prices, signal-impact drift remains < 4%. The next quarterly filing (Q2 2026, due ~2026-08-29) will automatically refresh SOXX even if iShares US stays blocked the whole quarter — SOXX dependency on a single operator's residential IP is now removed.
Strategy B · Asset-class momentum
The 12 ETFs are deliberately cap-weighted broad asset-class slates — US equity (SPY / IJR / QQQ), international developed (EFA / VGK / EWJ), real estate (VNQ), commodities (GLD / DBC), Treasuries (TLT / IEF / TIP). Within-equity rotation is Strategy A's job; B's role is the cross-asset switch. HYG was removed in 2026 — high-yield credit is equity-correlated when stress arrives and was never providing the “flight to safety” role its bond designation implied. EEM was moved to overlay-only in July 2026 (Phase 29) — the WS2 universe review found EEM double-counted between B's rotation and the Phase 22 tilt (look-through peaked at 15% of NAV); the role ablation showed all four role configurations within 0.009 Sharpe, so EM exposure is now expressed solely by the tilt and B is no worse without it (+1.02 vs +1.01 standalone). SHY (1–3y Treasury) is the cash floor for unfilled signal slots; was IEF before the May 2026 review, but IEF carries 7–10y duration risk in a rising-rate shock and is better held only when its own momentum signal earns the slot on its own merit. IEF remains in the rotation universe as a long-duration trade candidate.
Strategy C · Thematic optionality
The 25 themes were assembled through phased additions with explicit within-strategy and cross-strategy correlation gates: Tech & Innovation (ARKK, CIBR, SKYY, BOTZ, BLOK), Energy & Climate (ICLN, TAN, LIT, URA), Health & Bio (XBI, ARKG + IHI added 2026-05-31 for medical devices), Cyclical thematic (JETS), Commodity-equity (GDX, COPX, MOO + XME, WOOD, REMX added 2026-05), China & EM Tech (CQQQ added 2026-05, 159801.SZ added 2026-05.1 with FX-conversion pipeline), Infrastructure (PAVE + PHO added 2026-05-31 for water), Defence (ITA), Crypto (BTC-USD as the long-history proxy with IBIT-equivalent expense ratio applied; live execution in IBIT). May 2026 universe expansion: bulk screen of 27 candidates against the 23-theme universe — 5 passed Stage 1 (correlation < 0.85), 4 passed Stage 2 (walk-forward Sharpe degradation < 0.03 per candidate tested individually), 2 deployed (PHO + IHI). PRNT and BETZ passed both gates but were not deployed — too small to consistently make the K=4 cut. KRBN failed Stage 2. Deployment outcome honest disclosure: the screen tested each candidate alone; deploying BOTH together compounded the effect more than the per-candidate tests suggested. Strategy C standalone walk-forward Sharpe moved +0.52 → +0.45 (-0.07), but the deployed 35/35/10/20 blend Sharpe moved only -0.004 (within noise) — the 10% sleeve cap insulates the blend from C's standalone degradation. Net rationale for deployment: water-infrastructure (PHO) and medical-devices (IHI) fill structural exposure gaps next to PAVE and XBI; the cost is a small degradation in C's standalone metrics that does not propagate to the deployed portfolio. Two earlier candidates were tested and reverted despite passing the correlation gate: SLV (silver, dragged B's Sharpe by 0.18) and KWEB (China internet, dragged C's walk-forward Sharpe by 0.13). Lesson: correlation gates are necessary but not sufficient; empirical Sharpe / drawdown test is required before deployment. Honest caveat: severe survivorship bias — failed thematics (cannabis, leveraged-thematic, volatility ETFs) never entered the universe at all.
Strategy D · Europe sectors
The 5-ETF universe is structurally small because iShares Europe publishes a limited Stoxx Europe 600 sector UCITS slate. The five available with reliable constituent-CSV endpoints are Banks (EXV1), Oil & Gas (EXH1), Technology (EXV3), Industrial Goods & Services (EXH3), and Utilities (EXH9). Constituent breadth (the same signal mechanism as Strategy A) requires per-ETF holdings, which constrains the universe to issuers with reliable holdings-CSV endpoints. Single-country UCITS (IJPN Japan, NDIA India, ICHN China, ITWN Taiwan) were tested in Phase 4 as a “Countries” variant and deferred — insufficient single-country breadth data after the constituent-breadth pipeline fix exposed thin yfinance coverage for non-US tickers. The July 2026 WS2 review then evaluated the price-momentum version (10 country ETFs, top-K, own-200d gate) and rejected it: it beats an EEM+EFA benchmark only in EM-favoured regimes (3 of 6 sub-periods, train half negative) — a regime bet the Phase 22 tilt already expresses more cheaply. Future expansion (eg adding a 10-sector Europe set) would require iShares Europe to publish more sector UCITS or a switch to a different issuer with comparable holdings access.
Signal definition (formal)
Breadth indicator (Strategy A): % of an ETF's point-in-time constituents whose closing price is above their own 200-day simple moving average. Computed daily; constituents reset weekly from the iShares UK holdings endpoint.
Momentum signal (Strategy B): distance above own 200-day MA per ETF, (close − MA200) / MA200. Positive = uptrend, negative = downtrend. Computed daily from yfinance adjusted close.
Breadth indicator (Strategy D): same as Strategy A but on Stoxx Europe 600 sector UCITS constituents. (Pipeline note: min_periods in the rolling MA200 is relaxed from 100% to 90% of window so that non-US constituents with sparse missing days from local holidays or dividend events do not silently lose their MA. See the research log below for the bug-discovery story.)
Headline strategy — Multi-Strategy 35/35/10/20 A:B:C:D blend (no leverage):
- Strategy A: rank 14 US sector/broad ETFs by sector-relative breadth (each ETF's breadth minus the cross-sectional mean), take top K = 7, drop any below the mean (long-only), weight survivors by relative-breadth share. Holds up to 7 — fewer in narrow leadership regimes. Universe: SOXX, CSP1, CNDX, IUES, IUFS, IUHC, IUIS, IUCS, IUCD, IUUS, IUMS, IUCM, IUSP, IDP6. IUIT (S&P 500 Info Tech) was pruned in May 2026 because of 0.97 correlation with CNDX — see Robustness Test 12. Sector-relative signal deployed Phase 20.1 (May 2026) — lifted Sharpe +0.05 and cut SPY correlation 0.93 → 0.83 vs the prior absolute-breadth signal.
- Strategy B: rank 12 broad asset-class ETFs by distance above 200d MA, drop negatives, take top K = 7, weight by signal share, cash floor in SHY (1-3y Treasury) for unfilled slots. IEF remains in the rotation universe as a duration play. Universe (HYG removed 2026-05 — equity-correlated, never defensive; EEM moved to the Phase 22 overlay 2026-07, Phase 29): SPY, IJR, QQQ, EFA, VGK, EWJ, VNQ, GLD, DBC, TLT, IEF, TIP.
- Strategy C (10% sleeve): rank 25 thematic ETFs by distance above 200d MA, drop any below +5% (signal floor), take top K = 5, equal-weight (1/K) across the holdings, cash floor in SHY (1-3y Treasury) for unfilled slots. Sleeve-breadth gate (deployed June 2026): if fewer than 30% of the 25-theme universe clears the +5% floor at any rebal, exit ALL positions to SHY for the week. Universe: ARKK, CIBR, SKYY, BOTZ, BLOK, ICLN, TAN, LIT, URA, XBI, ARKG, JETS, GDX, COPX, MOO, PAVE, ITA (defence/aerospace), BTC-USD (Bitcoin spot — see caveat below; live execution in IBIT), XME (broad metals & mining), WOOD (timber & forestry), REMX (rare earth / strategic metals), CQQQ (Invesco China Technology — broad China-tech), 159801.SZ (Bosera CSI Chip ETF, CNY-denominated, USD-adjusted via spot FX — pure China A-share semiconductor hardware basket: Cambricon, AMEC, NAURA, SMIC. Deployed via IBKR Stock Connect; 588200.SS is an interchangeable Shanghai-listed alternative), PHO (Invesco Water Resources — water infrastructure, added 2026-05-31), IHI (iShares US Medical Devices — Abbott, Thermo Fisher, Medtronic, Stryker, added 2026-05-31). The C sleeve uses equal-weight (1/K) across the holdings rather than signal-weighting, after an A/B test showed equal-weight dominates on every metric for this high-floor (5%) regime — see the research log below.
- Strategy D (20% sleeve): rank 5 Stoxx Europe 600 sector UCITS by constituent breadth, take top K = 3, weight by breadth share. Universe: EXV1 (Banks), EXH1 (Oil & Gas), EXV3 (Technology), EXH3 (Industrial Goods & Services), EXH9 (Utilities). Trade as the Xetra-listed UCITS (EXV1.DE etc.) in EUR.
- Combination: 35% A + 35% B + 10% C + 20% D, rebalanced weekly Friday.
- Transaction cost: 5 bps per unit of weight change (10 bps round-trip).
- No look-ahead: every weight is decided using the prior trading day's signal.
- K is refit only annually on expanding-window Sharpe; choices have been stable at K = 7 (A), K = 7 (B), K = 5 (C, deployed June 2026 with sleeve-breadth gate), K = 3 (D).
Why 35/35/10/20 and not 45/45/10? The 4-way blend with the Europe sleeve posts the highest Sharpe of any variant tested (+1.15 vs +1.09 for the 3-way 45/45/10), at the cost of ~2.3pp wider max drawdown (-23.8% vs -21.5%). The Europe sleeve is real equity beta — it participates in 2020/2022 sell-offs and does not provide downside protection — but its idiosyncratic European macro cycle adds genuine diversification at the blend level. If drawdown floor is more important than Sharpe, stay with 45/45/10. If you want best risk-adjusted return, switch to 35/35/10/20.
Research log — major design decisions
Build-log entries that explain how the deployed configuration was chosen — architecture variants tested (adding Strategy D), universe-expansion attempts that failed the gate (sub-sector ETFs), the weighting-scheme A/B test (signal-weight vs equal-weight for Strategy C), the metric overhaul that surfaced statistical-significance limits, and the right-tail-metrics reframing of Strategy C as an optionality sleeve. All collapsed by default — the verdict and the current caveats are above, the reasoning is here.
Phase 4 — Strategy D + the constituent-breadth pipeline fix 2026-05-24
Phase 4 added Strategy D (Europe sector breadth) as a 4th sleeve. The build also forced a fix to the constituent breadth pipeline that quietly affected non-US ETFs.
What was tested: 9 new ETFs added to the registry — 5 Stoxx Europe 600 sector UCITS (EXV1, EXH1, EXV3, EXH3, EXH9) plus 4 single-country UCITS (IJPN, NDIA, ICHN, ITWN). The Phase 4 experiment ran 8 architecture variants: baseline 45/45/10 A:B:C, +Europe (heavy and light), +Countries (heavy and light), +both, MERGE (all into Strategy A's universe).
Result: Europe sleeve wins as a separate 20% sleeve (35/35/10/20 A:B:C:D, Sharpe +1.152 vs baseline +1.082 — delta +0.07). Countries lose on every variant (max Sharpe +0.99 vs baseline +1.08). MERGE loses (Sharpe +1.05). Countries deferred to a future phase pending universe expansion (only Japan + Taiwan had usable data after the pipeline issue was fixed; India + China constituent coverage remains thin in yfinance for non-US tickers).
Pipeline fix — the surprise discovery: while validating Strategy D output, noticed the equity curve was flat at 1.0 from 2018-01-26 to 2021-02-12 (3 years of "cash") then broke into trading, with breadth values frozen at the same number from 2023-04-06 onwards (3 years of "stale signal"). Investigation revealed:
compute_ma200_breadthinrun_ma200_sweep.pyusedrolling(200, min_periods=200).mean()— strict, requires ALL 200 observations in the window to be non-NaN.- US S&P 500 constituents have ~100% daily coverage via yfinance, so the strict requirement is a no-op for Strategy A.
- Non-US constituents (
.L,.DE,.PA,.AS,.MItickers) have 1-2% sparse missing days from local holidays and dividend events. Even constituents with 99% coverage failed the strict requirement because every 200-day window contained at least one missing day. Result:n_valid(count of constituents with computable MA200) collapsed to 0 → breadth went NaN →_build_panels_forffill'd the last valid value indefinitely. - Fix: relaxed
min_periodstoint(period * 0.9)(=180 for MA200) and tightened the denominator to require both today's price AND the MA to be valid. Now a constituent with at least 180 of the last 200 days valid contributes to the breadth signal — matching what a human trader would consider "enough history to call the trend". - Impact on US strategies: Strategy A Sharpe changed from +0.949 to +0.96 — well within rounding (US is the universe where the fix is a no-op). Strategy B and C use ETF-level momentum (not constituent breadth) so they were unaffected.
- Impact on Strategy D: was previously a broken +0.59 Sharpe with frozen post-2023 signal; now correctly +0.93 Sharpe across the honest 2018-2026 backtest with realistic -32% DD (was a fake -19% under the frozen-signal version).
Lesson: when extending a strategy mechanism to a new universe, the data plumbing assumptions may not hold. The bug had been latent the entire time we added IJPN / NDIA / ICHN / ITWN constituents — none of them produced usable breadth until the fix. Worth re-running the universe-level diagnostics on any future non-US additions to confirm continuous breadth.
Phase 5 — sub-sector universe expansion (negative result, reverted) 2026-05-24
Phase 5 tested 11 US sub-sector / industry ETFs (XME, AMLP, ITB, OIH, KRE, XRT, FDN, IBB, SMH, XOP, PBW, KIE, PHO, IGV) as candidates for expanding Strategy C's universe beyond the then-deployed 16 thematic ETFs (the live universe is now 23 — XME, WOOD, REMX, CQQQ, ITA, AIQ, IBIT, 159801.SZ were subsequently added in Phases 5–9). Most were SPDR / VanEck / First Trust / Invesco / Global X funds — none are iShares, so they would not fit Strategy A (which requires iShares holdings CSV), but the ETF-level momentum mechanism in Strategy C accommodates them with a one-line registry change per ticker.
Diagnostic gate (Test 12-style correlation analysis on weekly signal series):
- Within-Strategy-C gate (would the candidate just duplicate an existing C member?): 4 candidates failed — XME (cousin of COPX +0.87), IGV (cousin of SKYY +0.94), XRT (cousin of PAVE +0.86), FDN (cousin of SKYY +0.96). The user's initial screenshot intuition was wrong on these.
- Cross-strategy gate (would the candidate just amplify a sector Strategy A already holds?): 3 more candidates were blatant cousins of A's sector slate — XOP (XLE +0.95), OIH (XLE +0.91), KIE (XLF +0.91). Adding them would mostly just double-weight Energy / Financials when A is already there.
- Survivors: 4 candidates passed within-C but were marginal on cross-strategy — ITB (clean on both, max cross-A +0.75 with SPY), AMLP (XLE +0.85), PHO (SPY +0.85), KRE (XLF +0.87).
Empirical test: added the 4 survivors to Strategy C's UNIVERSE, re-ran Strategy C and the multi-strategy combinator. Results:
- Strategy C standalone Sharpe: +0.71 → +0.74 (+0.03 lift)
- Strategy C walk-forward Sharpe: +0.36 → +0.26 (-0.10 degradation — the deployment-quality metric got worse)
- Strategy C max drawdown: -43.7% → -41.6% (+2.1pp shallower)
- Deployed 4-way blend 35/35/10/20 Sharpe: +1.1503 → +1.1506 (+0.0003 — within noise)
- Deployed 4-way blend max DD: -23.8% → -23.8% (no change)
Verdict: reverted. The standalone Strategy C in-sample improvement is real but small; at the 10% sleeve weight in the deployed blend it is undetectable (+0.0003 Sharpe). The walk-forward degradation is the actual signal — the marginal cross-strategy cousins (AMLP / PHO / KRE riding XLE / SPY / XLF momentum cycles) chase fads that mean-revert out-of-sample. Combined with the operational cost of 4 extra holdings (AMLP issues K-1 tax forms which are annoying for SG/EU investors), no net benefit.
Lesson: the within-strategy correlation gate is necessary but not sufficient. The cross-strategy correlation gate matters too — at +0.85+ it kills the marginal additivity even when the within-strategy correlation is low. Future universe-expansion candidates should pass BOTH gates at threshold < 0.85 before being deployed. Sub-sector ETFs that are cousins of any sector already in Strategy A's slate (XLE, XLF, XLI, XLY, etc.) are particularly unproductive — Strategy A already captures that exposure efficiently.
What's queued for future research: (1) single-country breadth — needs universe expansion beyond the original 4 (consider adding EWZ Brazil, EWA Australia, EZA South Africa, EWS Singapore) and a more careful Strategy E architecture rather than treating countries as a sleeve of Strategy A. (2) MSCI ACWI ex-US sector breadth — if iShares offers UK UCITS that track non-US developed-market sectors at a finer granularity than the current Europe-only Strategy D, the same constituent-breadth mechanism could extend further.
Phase 6 — weighting-scheme A/B test (equal-weight shipped for C) 2026-05-24
Phase 6 tested the hypothesis that Strategy B and C's modest walk-forward Sharpe was partly caused by the signal-share weighting scheme, which heavily overweights the most-overbought ETF — statistically the one most likely to mean-revert.
The asymmetry: Strategy A uses breadth-share weighting on a bounded signal (breadth ∈ [0,1]) — top-1 weight typically ~16% of the sleeve. Strategy B and C use signal-share weighting on an unbounded signal (distance above 200d MA, can be +50% in bubbles) — top-1 weight can reach ~46% of the sleeve. That is 3× the concentration of A, on candidates that are statistically more reversal-prone.
Test: ran 4 weighting schemes side-by-side on Strategy C and Strategy B (no universe changes — pure parameter sweep):
- Current — signal-share (with 35% cap for C, no cap for B)
- Equal-weight — 1/K per holding (ignores signal magnitude)
- Sqrt(signal) — weight ∝ √signal (softens proportionality)
- Rank-weighted — top gets K weight units, K-th gets 1 (bounded dispersion regardless of signal magnitude)
Strategy C results: equal-weight dominates on every metric. IS Sharpe +0.708 → +0.781, WF Sharpe +0.364 → +0.388, CAGR +16.5% → +18.3%, max DD -43.7% → -42.8%, turnover 16.9× → 15.7×. The walk-forward K-sequence also shifts: the original scheme always picked the smallest K=3 (to compensate for the concentration the weighting creates); equal-weight alternates K=3 and K=5 — the rotation can hold more positions without giving any one too little weight.
Strategy B results: equal-weight has best IS Sharpe (+0.794 → +0.876) BUT max DD widens by 6pp (-14.6% → -20.9%). Sqrt(signal) is a softer alternative — best WF Sharpe (+0.743) with smaller DD damage (-17.2%). For B the signal-share scheme remains the best deployment choice. The DD trade-off matters: B's primary job is downside control via flight-to-bonds in crises, and that mechanism depends on weighting heavily into TLT/IEF when they are the strongest signals.
The mechanistic reason for the asymmetry: Strategy C has a +5% signal floor — eligible candidates are by definition already in well-established uptrends, so signal magnitude beyond eligibility carries little extra information (they're all strong). Weighting heavily toward the strongest just overweights the candidate most likely to mean-revert. Strategy B has 0% floor — eligible candidates include modest +0.5% above MA200, so signal magnitude does carry meaningful information about which trends are well-established vs. just starting. Signal-share is informative for B; equal-weight is informative for C.
Verdict: shipped equal-weight for Strategy C, kept signal-share for Strategy B and (already) breadth-share for A and D. The Strategy C standalone Sharpe lifts from +0.71 to +0.79; at the 10% sleeve weight, the 4-way deployed blend Sharpe lifts from +1.150 to +1.156. Modest in magnitude but consistent in direction with no downside (DD slightly improves, turnover slightly lower).
Generalisable lesson: the right weighting scheme depends on the signal-floor regime. With a high floor (C: +5%), the floor IS the filter — equal-weight after the floor. With a low floor (B: 0%), the magnitude IS the filter — signal-share with no floor. With a bounded signal (A and D: breadth ∈ [0,1]), signal-share is naturally tight and produces near-equal weights anyway, so the question does not arise.
Operational note: the original 35% per-ETF cap on Strategy C is now moot under equal-weight (K=5 → 20% per holding, well below the 35% cap). The cap code is retained in run_thematic_rotation.py as a no-op safeguard if K is ever reduced to 3 (where 33.3% per holding would still be under cap).
Phase 7 — metric overhaul + statistical-significance audit 2026-05-24
Phase 7 was prompted by a critique that the dashboard headline stats included three useless metrics (Total Return, Annual Turnover ×, Number of Rebalances) and missed the metrics an institutional AI would actually ask about. Three deliverables:
- Replaced headline stats across all four strategy tabs + the Multi-Strategy tab. Dropped Total Return (path-dependent, length-dependent, redundant with CAGR) and Number of Rebalances (carries zero information beyond what the cadence label already says). Added: Walk-forward Sharpe (the deployment-quality number, was buried in details for some sleeves), Calmar ratio (CAGR / |Max DD| — single best one-number risk-adjusted return), and average holding period in days (252 / annual_turnover — much more interpretable than "10.4×/yr"). Added an explicit cost drag sub-line (annual_turnover × 5 bps) so the dollar cost of execution is no longer hidden.
- Added walk-forward K refit to Strategy D, which previously had no OOS validation. Annual K refit on expanding train window, K ∈ {2, 3, 4}. Result: Strategy D walk-forward Sharpe is +0.97, HIGHER than in-sample +0.89. The K choice is stable (always K=3), the signal is genuinely persistent OOS, and the 6-segment K-sequence shows no overfitting on the cadence/K choice. This was the single biggest robustness gap in the Phase 4 deployment story; it now closes cleanly. WF Sharpe above IS Sharpe is an unusual and strong robustness result (most strategies see WF degradation from IS due to overfitting).
- Computed block-bootstrap CIs on every strategy's Sharpe and on the key paired differentials. Moving block bootstrap, block size 60 trading days (~3 months, matches Robustness Test 4 methodology), 2,000 samples, paired sampling preserves cross-strategy correlation. Output: per-strategy Sharpe with 95% CI, plus 4 paired differentials (deployed vs 3-way, deployed vs A alone, deployed vs 50/50 A:B, 3-way vs 50/50 A:B). Surfaced as a dedicated "Statistical significance" section on the Multi-Strategy tab and as a sub-line under each strategy tab's recipe-stats.
The honest finding from the bootstrap CIs:
- The 4-way deployed blend Sharpe (+1.16) is NOT statistically distinguishable from the 3-way baseline (+1.09) at 5% significance. The 95% CI on the differential is [-0.04, +0.17] — straddles zero. But the bootstrap gives the deployed blend a 83% probability of being a real improvement, which combined with the mechanistic rationale (Europe orthogonality, Phase 4 retrospective) is meaningful but not conclusive evidence.
- Strategy C's 10% sleeve contributes essentially zero at the blend level vs just 50/50 A:B. The 3-way 45/45/10 vs 2-way 50/50 differential is -0.004 Sharpe with p(better) = 41.9% — coin-flip. C earns its inclusion only as optionality on the next thematic bull run, not as a Sharpe lifter. This is now stated explicitly in the Multi-Strategy tab "How to choose your blend" section.
- The Phase 4-6 cumulative Sharpe improvement (+1.09 → +1.16, delta +0.06) is consistent with real improvement but within the noise floor of 7.5 years of data. Confidence will grow with more OOS years. The point estimates are all positive, the directions are consistent across multiple comparison points, and the mechanistic stories (Europe orthogonality for Phase 4, equal-weighting in high-floor regime for Phase 6) suggest the improvements are structurally motivated rather than data-snooped.
Implication for fund deployment: the deployed blend is the right point estimate, and the engineering process (correlation diagnostics, A/B tests, this bootstrap audit) is the right discipline. But the dashboard should not oversell the +0.06 Sharpe improvement as a definitive win — it is a directional improvement with strong mechanistic support, awaiting more OOS data for statistical confirmation. Anyone evaluating this as Navigo IP should understand both numbers: the point estimate and the CI.
What's still missing (Tier 2 backlog): live tracking infrastructure (paper-trade weekly, log actual-vs-backtest divergence), Bayesian / probabilistic Sharpe (richer than the existing bootstrap CI), regime-aware risk overlay (could reduce DD if it doesn't overfit), and tax-aware accounting documentation (K-1 forms for AMLP, European UCITS withholding, etc. — relevant for the SG investor base). Items already completed in Phase 11-12: realised cross-strategy return correlation matrix on Monitor (Phase 11A), hit rate + longest drawdown duration on each strategy tab (Phase 11B), per-strategy cost calibration (Phase 12 — see Caveats below).
Phase 8 — right-tail metrics and the optionality reframing for C 2026-05-24
Phase 8 was prompted by a substantive critique that the Phase 7 conclusion "Strategy C contributes essentially nothing at the blend level" was mean-variance-accurate but option-theoretically wrong. The point is generalisable: Sharpe ratio and bootstrap-on-Sharpe are the right metrics for symmetric alpha sleeves but the wrong metrics for asymmetric / optionality sleeves. Phase 8 builds the right metrics and reframes C correctly.
Why Sharpe underrates optionality strategies:
- Sharpe = mean ÷ standard deviation. Standard deviation treats positive and negative returns symmetrically.
- An optionality strategy is structured to be asymmetric: capped downside (Strategy C: 10% sleeve weight caps annual NAV impact at ~-10%), unbounded upside (if a thematic bull fires, C captures it; max sleeve impact ~+15% NAV per year empirically).
- The bootstrap-on-Sharpe p(better)=42% finding (Phase 7) is the mean outcome over 7.5 years. It implicitly assumes "the next 7.5 years look like the last 7.5 years". This assumption fails precisely when a new thematic regime emerges — the moment when C is most valuable.
- Backtests cannot price embedded options. The 2018-2026 window includes the 2020-21 thematic boom but cannot price the OPTIONS embedded in C for future booms we cannot name (space, quantum, longevity, fusion, AI-derivative themes).
What Phase 8 added — right-tail metrics (one new script + 5 new dashboard sections):
- Sortino ratio: annualised mean / annualised downside-only volatility. Credits upside vol; only penalises drawdowns. Fairer than Sharpe for convex strategies.
- Skewness of monthly returns: positive skew = right-tail bias.
- Best / worst rolling 12-month return + the specific date windows.
- Asymmetry ratio (|best 12m| / |worst 12m|): direct measure of right-tail dominance.
- % of months as top-performing sleeve: how often does each sleeve win? Captures the "when C wins, it wins often and big" property that bootstrap-on-Sharpe missed.
- Regime decomposition across 4 hand-picked sub-windows (Q4 2018 Powell pivot, COVID + thematic boom Mar 2020 → Feb 2021 ARKK peak, 2022 inflation crash, 2024 AI surge). Per-strategy total return + max DD in each.
Empirical findings — the C optionality case in numbers:
- C's best rolling 12-month return: +162%. Vs A +85% / B +43% / D +66%. C delivers the largest absolute upside tail in the universe by 2-4×.
- C is the top-performing sleeve 41% of months. A is top 21% / B is top 10% / D is top 28%. When C wins it wins more often than any other sleeve.
- COVID + thematic boom (Mar 2020 → Feb 2021): Strategy C standalone +170%. The 50/50 A:B blend (no C) got +58%. The 4-way blend (10% C) got +65%. The small C sleeve added ~+7pp to blend return in an 11-month window — capturing roughly 70% of C's idiosyncratic alpha at just 10% sleeve weight.
- 2022 inflation crash: C was worst (-25%), but the 10% cap limited blend damage to -13% (the 4-way blend was actually 2pp worse than the 50/50 A:B — a -2pp option premium for the +7pp option payoff). The asymmetric structure works.
- Q4 2018 Powell pivot: C +4.4% (defensive!) while A/B/D all -5 to -10%. The cash-floor mechanism (when fewer than K candidates clear the +5% floor, the deficit sat in IEF at the time; switched to SHY in mid-2026) kicked in during the sell-off.
The reframing: Strategy C earns its 10% sleeve weight through asymmetric convexity. Best 12m return is 2-4× any other strategy. Top sleeve 41% of months. The 10% sleeve cap structures it as a long-dated out-of-the-money call basket — small premium (the modest drawdown contribution), unbounded upside (any future thematic boom). Bootstrap-on-Sharpe gives p(better)=42% because Sharpe penalises upside vol; that finding is technically true but misses the optionality value entirely.
Generalisable lesson: when evaluating a strategy or sleeve, ask "is this symmetric alpha or asymmetric optionality?" first. Symmetric alpha → Sharpe + bootstrap CI is the right gate. Asymmetric optionality → right-tail metrics + regime decomposition + sleeve-cap structure is the right gate. Using the wrong gate leads to wrong conclusions; Phase 7 nearly did so for C.
Implication for deployment: the 4-way 35/35/10/20 blend remains the deployed default. Phase 8 strengthens the case for keeping the 10% C sleeve rather than dropping it to deploy 50/50 A:B. The honest framing is: "core risk-adjusted return from A+B+D (~+1.13 Sharpe pre-overlay; +1.31 deployed), with a 10% optionality sleeve via C that delivered +170% during the 2020-21 thematic boom and is positioned to capture the next regime that emerges". The Sharpe-only framing undersells the structure.
Universe and data sources
14 ETFs in the universe: SOXX (semis), CSP1 (S&P 500), CNDX (NASDAQ-100), and the S&P 500 sector slate — IUES (Energy), IUFS (Financials), IUHC (Health Care), IUIS (Industrials), IUCS (Consumer Staples), IUCD (Consumer Discretionary), IUUS (Utilities), IUMS (Materials), IUCM (Communication Services). Plus IUSP (US REITs as Real Estate proxy — see caveat) and IDP6 (S&P SmallCap 600).
Pruned (May 2026): IUIT (S&P 500 Info Tech) was removed because of 0.97 correlation with CNDX (NASDAQ-100). CNDX has much higher trading liquidity via its US-listed equivalent QQQ, so keeping IUIT was double-counting the same large-cap tech bet with a less-tradeable variant. The prune improved walk-forward Sharpe by +0.034 (Robustness Test 12).
Trading proxies: SPDR Select Sector ETFs (XLE, XLF, XLV, XLI, XLP, XLY, XLU, XLB, XLC, XLRE). CSP1 trades via SPY; CNDX via QQQ; SOXX trades itself; IDP6 trades via IJR.
Real Estate caveat: no iShares UK S&P 500 Real Estate sector UCITS exists as of 2026-05. IUSP (FTSE EPRA NAREIT US Dividend+ Index, ~38 US REITs) is used as a substitute. The constituent set is broader than the S&P 500 Real Estate sub-sector, but the breadth metric (% above 200d MA) is still constituent-relative and remains comparable.
Data sources: iShares UK holdings endpoint for point-in-time constituents (weekly Friday snapshots back to 2018-01-05), yfinance for adjusted close prices of constituents and trading proxies. All data is free and reproducible from the scripts in this repo.
OOS validation
Both Strategy A and Strategy B are validated walk-forward (see Robustness → Test 10 for Strategy A's annual K refit and Strategy B's tab for its own walk-forward). The breadth signal itself was earlier validated via a 2022-09-08 train/test split on a composite breadth indicator: the train-half winner held up on the test half (train Sharpe +0.94 → test Sharpe +1.42 on SOXX; train winner ranked #1 of 11 cells on test grid). The current top-K rotation paradigm strips away the per-ETF threshold tuning that earlier work relied on, eliminating most of the in-sample selection bias.
Best practices for parameter optimisation (lessons from this project)
The Robustness tab demonstrates that the headline Sharpe was materially inflated by in-sample optimisation. Nine rules of thumb learned the hard way:
- Always report walk-forward (or train/test) Sharpe alongside in-sample. The in-sample number is biased upward by ~0.3 Sharpe in our case. If you only have IS, halve your conviction.
- Robust over optimal. A parameter that's "second-best" but stable across sub-windows is more deployable than a parameter that's "best" but jumps each refit. Our walk-forward L sequence for CSP1 was [55, 80, 75, 75, 75] — wildly unstable; deploying any single one of those would have given different results.
- Reduce degrees of freedom. Per-ETF tuning adds N free parameters. A single global parameter (e.g., L=60 for all 11 ETFs at the time of the Test 10 robustness study; the live universe is now 14) had ZERO tuning cost, and beat walk-forward L on 8 of 11 ETFs. Less is more.
- Use external priors. If literature gives you a number (Zweig's L=60 breadth threshold), USE IT. A prior-based value has no overfit cost; a data-driven value does.
- Lower the rebalance frequency to lower in-sample noise transmission. Daily rebalancing transmits each fitted-L noise into more daily decisions. Bi-weekly rebalancing reduces the effective number of "actions" per unit time, smoothing the noise.
- Stress-test by regime, not just full-window Sharpe. The sub-period decomposition showed our strategy underperformed BH in 2019 pre-COVID, 2022 inflation shock, and 2023 AI rally. Full-window Sharpe hid this. Decompose by regime — if performance is concentrated in one period, you're sampling-lucky.
- Avoid the joint maximum across multiple parameters. If you can choose L, MA-period, cadence, AND base/thrust allocation, picking the joint argmax has compounded selection bias. Pick one parameter at a time on independent data, or use a fixed-parameter version.
- If walk-forward Sharpe ≤ BH Sharpe, abandon the strategy. No amount of in-sample tweaking saves a strategy whose held-out performance is no better than the passive benchmark. CSP1 fails this test (WF 0.57 vs BH 0.84); SOXX with bi-weekly cadence passes (1.05 vs 0.98).
- Prefer relative signals to absolute thresholds. "Top K by breadth" is a relative statement that self-normalises across regimes; "breadth ≥ L" is an absolute threshold that has to be re-tuned every regime. Test 10 shows the cross-sectional rotation paradigm has zero walk-forward degradation, while threshold-based timing degrades ~0.3 Sharpe. Whenever a signal can be expressed as a cross-sectional rank instead of a level, prefer the rank version.
Practical deployment translation. The strategy that survives all the tests is paradigm 3 from Test 10: own up to K = 5-7 ETFs by current MA200 breadth, ranked on sector-relative breadth (breadth minus the cross-sectional mean), drop sectors below the mean to keep the sleeve long-only, weight survivors by their positive relative-breadth share, rebalance weekly Friday, no per-ETF threshold tuning, refit K only annually. Walk-forward Sharpe ≈ 1.05 with no degradation from in-sample. The previous "fixed L=60 on IUIT / CNDX / IUES" recommendation was correct given the framing of paradigm 2, but paradigm 3 is strictly better — it owns the same sectors when they lead, drops them when they lag, and never bets on a single ETF being "above the L-line" in isolation. The original headline Sharpe 1.04 for CSP1 single-ETF was an artifact; the headline Sharpe 1.05 for top-K rotation is honest.
Caveats — what this backtest does not capture
Honest limitations of the deployed strategy as documented:
- Slippage on rebalances: rebalances assume mid-spread fills. Worst-case slippage may add another 5-10 bps per rebalance on top of the 5 bps transaction cost already modelled.
- Per-sleeve cost calibration: replaced the uniform 5 bps with sleeve-specific costs that reflect the actual liquidity of each universe. Strategy A = 2 bps (very liquid SPDR Select Sector + SPY/QQQ). Strategy B = 2 bps (SPY/IEF/GLD/TLT — among the most liquid ETFs in the world). Strategy C = 5 bps (mid-liquidity thematics, mixed: ARKK/XBI are 1-3 bps but BLOK/PAVE/BOTZ are 5-10 bps). Strategy D = 9 bps (European UCITS on Xetra at 5-10 bps bid-ask plus 2-4 bps FX cost for USD-base investors). Net effect: deployed blend Sharpe lifted from +1.16 to +1.18 (~+0.015), with A and B benefiting from tighter realistic costs more than D loses from wider European spreads.
- Survivorship in some price series: a few historically-acquired constituents (~5-15% of the early-window roster for some ETFs) are missing from yfinance, biasing early-window breadth measurements toward survivors. Less of an issue post-2020. A pipeline fix to
compute_ma200_breadthresolved the worse latent issue for non-US constituents (see research log below). - Common window for the deployed blend is 2018-11 → 2026-05 (~7.5 years), constrained by Strategy C's data start date. Strategy B's 18-year history (2008-2026, covering GFC + COVID + 2022) cushions the blend through real crises in the per-strategy comparison, but the blend-level statistics only see 2020/2022.
- Concentration risk in Strategy A: even at K = 7 the top breadth tilts toward 1-2 dominant sectors, so a semis or energy reversal can move 10-20% of NAV within the sleeve in a week. The 35% sleeve weight in the 4-way blend caps blend-level exposure but does not eliminate this.
- Strategy B cash-floor mechanic: when fewer than K asset classes are positive-trend, the deficit sits in SHY (1-3y Treasury, duration-neutral). Earns short-end Treasury carry with minimal duration risk; IEF (7-10y) stays in the rotation universe as a long-duration play but is no longer the cash proxy.
- Strategy C walk-forward Sharpe is +0.51 vs in-sample +0.82. The recent commodity-equity additions (XME, WOOD, REMX) plus CQQQ are net-neutral on in-sample (−0.01) and add a small drag on walk-forward (−0.04 versus the prior universe) — but the blend-level cost is essentially zero (−0.002 Sharpe) while max DD improves by 0.27pp. The 10% sleeve cap is the risk-management response to expected thematic-momentum degradation. C earns its place on right-tail / regime metrics and on cumulative blend-level improvements as the universe expanded.
- China A-share semiconductor exposure (159801.SZ) added via FX-aware download pipeline. 159801.SZ is the Bosera CSI Chip ETF, CNY-denominated and Shenzhen-listed. The backtest pipeline now downloads CNY prices, applies a daily FX conversion using yfinance's USDCNY=X spot, and reindexes onto the NYSE trading calendar with a 10-day stale-fill cap (covers Chinese New Year + October Golden Week SSE/SZSE closures). A 50bps annual expense ratio is applied as per-calendar-day compounded drag, same pattern as IBIT's 25bps on BTC-USD. Live execution is via IBKR Stock Connect; 588200.SS (Harvest SSE STAR Chip) is an interchangeable Shanghai-listed alternative with 0.96 correlation to 159801.SZ but insufficient history (3.65y vs the 5y walk-forward minimum). Also: because 159801.SZ inception (2019-08) is after BLOK's 2018-01 binding date, the script's eligibility logic was refactored to treat it as a late-inception ticker that does not constrain the backtest window start — without this fix, naively adding 159801.SZ would have collapsed the backtest from 7.5y to ~6y. The momentum picker (currently K=5, was K=4 pre-V6) excludes NaN signals automatically, so the late-inception decoupling is safe and lossless.
- SLV (B) and KWEB (C) were tested and reverted. Two illustrative cases of correlation-gate-passing candidates that empirically failed the deployment test:
- SLV (silver) passed Strategy B's gate (0.78 vs GLD) but dragged B's Sharpe from +0.99 to +0.81 (−0.18) and widened max DD from −14% to −27% (+13pp worse). Mechanism: silver's chop-then-reverse profile poisons a top-K momentum signal in a way GLD's smoother trend behaviour does not.
- KWEB (China Internet) passed Strategy C's gate (0.57 vs LIT) but together with CQQQ dragged Strategy C's walk-forward Sharpe from +0.50 to +0.37 (−0.13). Mechanism: the 2021-2023 China internet crackdown (−70% over two years) created a unique drawdown profile that the momentum signal (then K=4, now K=5 + sleeve-breadth gate) repeatedly bounce-traded and got chopped on. CQQQ-alone keeps the China-tech diversification without KWEB's crackdown-specific drag.
- Bitcoin (BTC-USD) is included via a spot-index proxy with IBIT-equivalent cost. IBIT (iShares Bitcoin Trust) only launched 2024-01-11, too short for the 5-year walk-forward methodology. The strategy therefore backtests using the CoinDesk Bitcoin Reference Price (BTC-USD, 8.4y of history) with IBIT's 25 bps annual expense ratio applied as a per-calendar-day compounded price drag, so the historical return path matches what an IBIT holder would have paid. Live execution remains in IBIT, which tracks BTC-USD with negligible error post-launch. GBTC (Grayscale Bitcoin Trust) was considered as the backfill but rejected: its market price traded at a 10-50% premium to NAV 2015-2021 and a 20-50% discount 2022-2023 (collapsing only on the Jan 2024 ETF conversion), so GBTC-price momentum would have captured the GBTC-discount narrative rather than BTC momentum. BTC-USD has none of that fund-structure noise. Per-ETF cap (PER_ETF_CAP = 35%, but equal-weight allocator gives 1/K = 25%) × 10% sleeve weight = max ~2.5% of NAV in BTC at any one time.
- Strategy D constituent breadth depends on a pipeline fix:
compute_ma200_breadthoriginally usedmin_periods=200which silently dropped the MA200 for any non-US constituent with sparse missing days. Without the fix (relaxed to 90% of window), D's signal froze on a 2023-04-06 ffill'd value. Documented in the research log below. - The +0.06 deployed-vs-baseline Sharpe improvement is not statistically significant at 5% (block-bootstrap 95% CI [-0.04, +0.17], p(better) 83%). Directionally supported by mechanistic rationale + consistent across multiple comparison points, but await more OOS data for statistical confirmation.
Repo + reproduce
Full source at github.com/phuazz/breadth-thrust-etf. To reproduce the deployed dashboard end-to-end:
Data fetching (one-off per ETF):
python scripts/fetch_constituents.py --etf {SOXX | CSP1 | CNDX | IUES | ... | EXV1 | EXH1 | ... }for each of the 14 US sector + 5 Europe sector ETF symbols (seescripts/etf_registry.pyfor the full list)python scripts/compute_breadth.py --etf {symbol}for each ETF that uses constituent breadth (Strategy A + D universes)
Strategy engines (per-deployment refresh):
python scripts/run_topk_robustness.py— Strategy A: top-K rotation, K × cadence grid, walk-forward, trade historypython scripts/run_asset_class_rotation.py— Strategy B: asset-class momentum rotation (13 broad ETFs, HYG removed)python scripts/run_thematic_rotation.py— Strategy C: thematic momentum rotation (23 ETFs, equal-weight)python scripts/run_europe_rotation.py— Strategy D: Europe sector breadth rotation (5 Stoxx Europe 600 sector UCITS), with walk-forward K refitpython scripts/run_multi_strategy.py— combines A+B+C+D into 14 blend variants (3 4-way + 3 3-way + 3 2-way + 4 standalones + meta-rotation)
Validation & audit:
python scripts/run_phase7_bootstrap.py— moving block bootstrap on per-strategy + paired-differential Sharpe CIspython scripts/run_phase8_right_tail.py— Sortino, skewness, rolling 12m extremes, regime decomposition, % months as top sleeve
Dashboard build:
python scripts/pipeline.py— injects all JSONs intotemplate.html→ writesdocs/index.htmlfor GitHub Pages
Other:
python scripts/run_ma200_sweep.py— per-ETF MA200 baselines (legacy: feeds the Monitor tab's per-ETF state cards)python scripts/run_robustness.py— runs the historical robustness suite (Tests 10/11/12 plus legacy Tests 1-9 which are no longer rendered in the dashboard but remain in the JSON for archival reproducibility)
An earlier iteration explored a more complex composite breadth signal (RSI breadth + MA breadth + Highs breadth + thrust detection) before the simpler MA200 sweep showed it generalised worse. That work is in the git history. Earlier work also tested per-ETF L-threshold strategies (legacy Robustness Tests 1-9) before top-K rotation was chosen as the deployed paradigm.
Data Health
Every data feed the dashboard reads, with freshness status as of the last pipeline build. A red row means the signal panels that read from it could be stale — do not deploy off that signal until the row turns green. Click fix command on a red row to see exactly which script to run.
| Status | Feed | Group | Last data | Days old | Fix command |
|---|
Thresholds: live feeds warn at 3d, stale at 7d. Strategy / blend / overlay outputs warn at 8d, stale at 14d (weekly cadence). Per-ETF breadth panels warn at 8d, stale at 14d. Constituent rosters warn at 14d, stale at 60d. SOXX uses wider thresholds (breadth 21d/60d, roster 30d/90d) because the iShares-US holdings endpoint is Akamai-blocked from most IPs and the script automatically carries the most recent good roster forward.
Fix: Most cases are resolved by python scripts/refresh_all.py from the repo root. Individual scripts shown per row let you target a single feed.