Files
ftdt-quant-lab/notebooks/02_strategy_research.ipynb
T
ramseshk 0446443d36 feat: creative alpha models + portfolio layer targeting Sharpe > 1.5
New strategies:
  - Cross-Sectional Momentum: long top-N, short bottom-N across HL universe
  - Spot-Perp Basis Arbitrage: delta-neutral spot vs perp price gap trading
  - Regime-Switching Ensemble: dynamically allocates strategies by market regime
  - Portfolio Construction: risk parity, vol targeting, correlation penalty

Infrastructure:
  - DuckDBDataProvider: real tick/candle data for backtests (replaces synthetic)
  - Walk-Forward Validation: systematic IS/OOS across all 12 strategies
  - 3 Jupyter research notebooks (EDA, strategy research, portfolio)

Pipeline integration:
  - deploy.py registry, sweep_runner, vbt_runner all updated
  - fee_tiers support for new strategies
  - All modules syntax-validated and import-tested
2026-08-12 12:26:29 +08:00

428 lines
16 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"cells": [
{
"cell_type": "markdown",
"id": "5aa80ede",
"metadata": {},
"source": [
"# FTDT Quant Lab — Strategy Research & Backtesting\n",
"\n",
"**Goal:** Develop and validate systematic trading strategies targeting Sharpe > 1.5 on Hyperliquid assets.\n",
"\n",
"**Framework:**\n",
"1. Signal Generation — compute alpha from market data\n",
"2. VectorBT Backtest — fast vectorized simulation with fee-accurate PnL\n",
"3. Walk-Forward Validation — IS/OOS parameter optimization\n",
"4. Statistical Significance — DSR, PSR, Sharpe Haircut, QuantVerdict\n",
"5. Deployment Decision — DEPLOY / SIMULATE / DISCARD\n",
"\n",
"**Key Thresholds for Sharpe > 1.5:**\n",
"- Win rate > 55% with positive expectancy\n",
"- Max drawdown < 15%\n",
"- Walk-forward consistency > 60%\n",
"- DSR > 0.80, PSR > 0.70\n",
"- Average trade PnL > 2x fees"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7e17950c",
"metadata": {},
"outputs": [],
"source": [
"# Setup\n",
"import sys; sys.path.insert(0, str(Path.cwd().parent))\n",
"\n",
"import numpy as np\n",
"import pandas as pd\n",
"import matplotlib.pyplot as plt\n",
"import seaborn as sns\n",
"from pathlib import Path\n",
"import json, time\n",
"\n",
"from backtests.vbt_runner import VBTBacktestRunner\n",
"from backtests.vbt_validator import VBTValidator\n",
"from quant.significance import QuantVerdict, validate_strategy\n",
"from quant.walkforward import WalkForwardRunner, quick_validate\n",
"from quant.optimizer import ParamOptimizer\n",
"from framework.data import HyperliquidDataProvider\n",
"from config.fee_tiers import get_perp_fees, get_strategy_fee_model\n",
"\n",
"sns.set_theme(style=\"darkgrid\")\n",
"plt.rcParams[\"figure.figsize\"] = (14, 6)\n"
]
},
{
"cell_type": "markdown",
"id": "ac92b47f",
"metadata": {},
"source": [
"## 1. Strategy Inventory\n",
"\n",
"Current strategies and their signal logic:"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "bd9d0ae0",
"metadata": {},
"outputs": [],
"source": [
"strategies = {\n",
" \"pairs\": {\n",
" \"name\": \"Pairs Trading\",\n",
" \"type\": \"Stat Arb\",\n",
" \"signal\": \"BTC/ETH ratio Z-score\",\n",
" \"entry\": \"|Z| > 1.5σ\",\n",
" \"exit\": \"|Z| < 0.5σ\",\n",
" \"best_use\": \"Range-bound, mean-reverting markets\",\n",
" \"sharpe_target\": 2.0,\n",
" },\n",
" \"hurst_vpin\": {\n",
" \"name\": \"Hurst VPIN\",\n",
" \"type\": \"Directional\",\n",
" \"signal\": \"Hurst > 0.55 AND VPIN > 0.25\",\n",
" \"entry\": \"Both trending + high flow imbalance\",\n",
" \"exit\": \"Hurst < 0.45 or direction flip\",\n",
" \"best_use\": \"Trending, high-volume markets\",\n",
" \"sharpe_target\": 2.5,\n",
" },\n",
" \"cross_sectional\": {\n",
" \"name\": \"Cross-Sectional Momentum\",\n",
" \"type\": \"Multi-Asset Long/Short\",\n",
" \"signal\": \"Past N-bar return ranking\",\n",
" \"entry\": \"Long top-3, short bottom-3\",\n",
" \"exit\": \"Next rebalance period\",\n",
" \"best_use\": \"All regimes, best in TRENDING\",\n",
" \"sharpe_target\": 1.8,\n",
" },\n",
" \"spot_perp_basis\": {\n",
" \"name\": \"Spot-Perp Basis Arb\",\n",
" \"type\": \"Delta-Neutral Carry\",\n",
" \"signal\": \"Perp vs spot price gap > 3bps\",\n",
" \"entry\": \"Short premium leg, long discount leg\",\n",
" \"exit\": \"Basis convergence < 1bps\",\n",
" \"best_use\": \"FUNDING_EXTREME, volatile basis\",\n",
" \"sharpe_target\": 2.0,\n",
" },\n",
" \"regime_ensemble\": {\n",
" \"name\": \"Regime-Switching Ensemble\",\n",
" \"type\": \"Meta-Strategy\",\n",
" \"signal\": \"Regime × strategy affinity matrix\",\n",
" \"entry\": \"Weights strategies by regime fit\",\n",
" \"exit\": \"Regime change or signal fade\",\n",
" \"best_use\": \"All environments — adapts dynamically\",\n",
" \"sharpe_target\": 2.0,\n",
" },\n",
" \"grid_mm\": {\n",
" \"name\": \"Grid Market Making\",\n",
" \"type\": \"Market Making\",\n",
" \"signal\": \"Symmetric grid around mid\",\n",
" \"entry\": \"Grid fill triggers position\",\n",
" \"exit\": \"Grid exit on rebalance\",\n",
" \"best_use\": \"LOW_VOL, CHOPPY\",\n",
" \"sharpe_target\": 2.0,\n",
" },\n",
" \"as_mm\": {\n",
" \"name\": \"Avellaneda-Stoikov MM\",\n",
" \"type\": \"Market Making\",\n",
" \"signal\": \"Reservation price from inventory\",\n",
" \"entry\": \"Reservation > best bid (buy) / < best ask (sell)\",\n",
" \"exit\": \"Hold period or profit target\",\n",
" \"best_use\": \"LOW_VOL with tight spreads\",\n",
" \"sharpe_target\": 1.5,\n",
" },\n",
"}\n",
"\n",
"for key, s in strategies.items():\n",
" print(f\"\\n{s['name']} ({key})\")\n",
" print(f\" Type: {s['type']}\")\n",
" print(f\" Signal: {s['signal']}\")\n",
" print(f\" Entry: {s['entry']}\")\n",
" print(f\" Exit: {s['exit']}\")\n",
" print(f\" Regime: {s['best_use']}\")\n",
" print(f\" Target Sharpe: {s['sharpe_target']}\")\n"
]
},
{
"cell_type": "markdown",
"id": "3f75e8a6",
"metadata": {},
"source": [
"## 2. Backtest Harness\n",
"\n",
"Run any strategy through the VBT backtest engine with fee-accurate PnL, then\n",
"validate with statistical significance tests."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "78c6a707",
"metadata": {},
"outputs": [],
"source": [
"def run_and_validate(strategy, interval='1h', params=None):\n",
" '''Run a full backtest + statistical validation pipeline.'''\n",
" print(f\"\\n{'='*60}\")\n",
" print(f\" {strategies.get(strategy, {}).get('name', strategy)} — {interval}\")\n",
" print(f\"{'='*60}\")\n",
" \n",
" runner = VBTBacktestRunner(vip_tier=0, staking_tier='none')\n",
" result = runner.run_strategy(\n",
" strategy=strategy, interval=interval, testnet=False, limit=500, params=params\n",
" )\n",
" \n",
" if result is None:\n",
" print(f\" No result (no trades or data error)\")\n",
" return None\n",
" \n",
" # Display key metrics\n",
" print(f\" Sharpe: {result.get('sharpe', 0):.3f}\")\n",
" print(f\" Total Return: {result.get('total_return_pct', 0):.1f}%\")\n",
" print(f\" Max Drawdown: {result.get('max_drawdown_pct', 0):.1f}%\")\n",
" print(f\" Win Rate: {result.get('win_rate', 0)*100:.0f}%\")\n",
" print(f\" Profit Factor: {result.get('profit_factor', 0):.2f}\")\n",
" print(f\" Total Trades: {result.get('total_trades', 0)}\")\n",
" print(f\" PnL: ${result.get('pnl', 0):.2f}\")\n",
" \n",
" # Statistical validation\n",
" n_trades = max(result.get('total_trades', 1), 1)\n",
" verdict = validate_strategy(\n",
" sharpe=result.get('sharpe', 0),\n",
" n_trades=n_trades,\n",
" n_trials=10,\n",
" wf_consistency=0.7,\n",
" )\n",
" print(f\"\\n Verdict: {verdict['verdict']}\")\n",
" print(f\" DSR (deflated): {verdict['deflated_sharpe']:.3f}\")\n",
" print(f\" PSR: {verdict['psr']:.3f}\")\n",
" print(f\" Haircut Sharpe: {verdict['haircut_sharpe']:.3f}\")\n",
" print(f\" Score: {verdict['score']}\")\n",
" print(f\" → {verdict['recommendation']}\")\n",
" \n",
" # Plot equity curve\n",
" eq = result.get('equity_curve')\n",
" if eq:\n",
" df_eq = pd.DataFrame(eq)\n",
" df_eq['t'] = pd.to_datetime(df_eq['t'])\n",
" df_eq.set_index('t', inplace=True)\n",
" df_eq['v'].plot()\n",
" plt.title(f\"{strategies.get(strategy, {}).get('name', strategy)} — Equity Curve\")\n",
" plt.ylabel('Equity ($)')\n",
" plt.show()\n",
" \n",
" return result\n",
"\n",
"# Quick sweep of top strategies\n",
"for s in [\"pairs\", \"hurst_vpin\", \"grid_mm\", \"momentum\", \"mean_rev\"]:\n",
" run_and_validate(s, \"1h\")\n"
]
},
{
"cell_type": "markdown",
"id": "eba6e554",
"metadata": {},
"source": [
"## 3. Cross-Sectional Momentum Backtest\n",
"\n",
"The new multi-asset strategy. Long the top performers, short the laggards."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "76c83fb1",
"metadata": {},
"outputs": [],
"source": [
"from strategies.cross_sectional_momentum import CrossSectionalMomentum, HIGH_LIQUIDITY\n",
"from data.duckdb_provider import DuckDBProvider\n",
"\n",
"duckdb = DuckDBProvider()\n",
"\n",
"# Fetch multi-asset candles\n",
"coins = [\"BTC\", \"ETH\", \"SOL\", \"HYPE\", \"ARB\", \"OP\"]\n",
"prices = duckdb.fetch_multi_candles(coins, interval='1h', limit=500)\n",
"\n",
"print(f\"Coins with data: {list(prices.keys())}\")\n",
"for coin in sorted(prices):\n",
" df = prices[coin]\n",
" print(f\" {coin}: {len(df)} bars, close=${df['close'].iloc[-1]:.2f}\")\n",
"\n",
"# Compute cross-sectional momentum signals\n",
"cs_mom = CrossSectionalMomentum(lookback=20, top_n=2, bottom_n=2, risk_parity=True, vol_target=0.20)\n",
"close_prices = {c: df['close'] for c, df in prices.items()}\n",
"weights = cs_mom.compute_signals(close_prices)\n",
"\n",
"print(f\"\\nCross-Sectional Momentum Weights:\")\n",
"for coin, wt in sorted(weights.items(), key=lambda x: abs(x[1]), reverse=True):\n",
" direction = \"LONG\" if wt > 0 else \"SHORT\"\n",
" print(f\" {coin:6s}: {direction:5s} {wt:+.3f}\")\n"
]
},
{
"cell_type": "markdown",
"id": "37dbc044",
"metadata": {},
"source": [
"## 4. Walk-Forward Parameter Optimization\n",
"\n",
"For strategies that show promise, run walk-forward to find stable parameters\n",
"and validate OOS performance."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "b264d513",
"metadata": {},
"outputs": [],
"source": [
"from quant.optimizer import ParamOptimizer\n",
"\n",
"# Grid MM parameter sweep\n",
"print(\"=== Grid Market Making — Parameter Optimization ===\\n\")\n",
"\n",
"opt = ParamOptimizer(strategy='grid_mm', interval='1h', coin='BTC', n_windows=3)\n",
"opt.add_param('grid_levels', [5, 10, 20])\n",
"opt.add_param('spacing_bps', [2, 5, 10])\n",
"opt.add_param('rebalance_every', [5, 10, 20])\n",
"\n",
"optimizer = ParamOptimizer.__new__(ParamOptimizer)\n",
"# [MANUAL RUN REQUIRED — uses live HL API, uncomment to run]\n",
"# report = opt.run()\n",
"# report.print()\n",
"print(\" Walk-forward optimizer ready. Uncomment `opt.run()` to execute (requires live HL API data).\")\n",
"print(\" Grid: 3 grid_levels × 3 spacing × 3 rebalance = 27 combinations × 3 windows = 81 backtests\")\n"
]
},
{
"cell_type": "markdown",
"id": "0c623f81",
"metadata": {},
"source": [
"## 5. Pairs Trading Deep Dive\n",
"\n",
"The only live-profitable strategy. Analyze its performance characteristics\n",
"and identify improvement opportunities."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "44344f2c",
"metadata": {},
"outputs": [],
"source": [
"# Pairs trading: analyze BTC/ETH spread dynamics\n",
"btc = prices.get('BTC', {}).get('close')\n",
"eth = prices.get('ETH', {}).get('close')\n",
"\n",
"if btc is not None and eth is not None and not btc.empty and not eth.empty:\n",
" common_idx = btc.index.intersection(eth.index)\n",
" btc = btc[common_idx]\n",
" eth = eth[common_idx]\n",
" \n",
" ratio = btc / eth\n",
" mu = ratio.rolling(20).mean()\n",
" std = ratio.rolling(20).std()\n",
" z_score = (ratio - mu) / std\n",
" \n",
" fig, (ax1, ax2, ax3) = plt.subplots(3, 1, figsize=(16, 12), sharex=True)\n",
" \n",
" ax1.plot(ratio.index, ratio, linewidth=0.5, color='black', label='BTC/ETH Ratio')\n",
" ax1.plot(mu.index, mu, linewidth=1, color='blue', label='20-bar MA')\n",
" ax1.fill_between(mu.index, mu - 2*std, mu + 2*std, alpha=0.15, color='blue', label='±2σ')\n",
" ax1.legend()\n",
" ax1.set_title('BTC/ETH Ratio with Bollinger Bands')\n",
" \n",
" ax2.plot(z_score.index, z_score, linewidth=0.5, color='purple')\n",
" ax2.axhline(1.5, color='red', linestyle='--', alpha=0.5, label='Entry (1.5σ)')\n",
" ax2.axhline(-1.5, color='red', linestyle='--', alpha=0.5)\n",
" ax2.axhline(0.5, color='green', linestyle='--', alpha=0.3, label='Exit (0.5σ)')\n",
" ax2.axhline(-0.5, color='green', linestyle='--', alpha=0.3)\n",
" ax2.legend()\n",
" ax2.set_ylabel('Z-Score')\n",
" \n",
" ax3.plot(z_score.index, abs(z_score), linewidth=0.5, color='orange')\n",
" ax3.axhline(1.5, color='red', linestyle='--', alpha=0.5)\n",
" ax3.set_ylabel('|Z|')\n",
" ax3.set_xlabel('Date')\n",
" \n",
" plt.tight_layout()\n",
" plt.show()\n",
" \n",
" # Signal statistics\n",
" entry_count = (abs(z_score) > 1.5).sum()\n",
" exit_count = ((abs(z_score.shift(1)) > 0.5) & (abs(z_score) < 0.5)).sum()\n",
" print(f\"Entry signals (|Z| > 1.5): {entry_count}\")\n",
" print(f\"Exit signals (|Z| < 0.5): {exit_count}\")\n",
" print(f\"Signal density: {entry_count / len(z_score) * 100:.1f}%\")\n",
" \n",
" # Distribution of Z-scores\n",
" print(f\"\\nZ-Score distribution:\")\n",
" print(f\" Mean: {z_score.mean():.3f}\")\n",
" print(f\" Std: {z_score.std():.3f}\")\n",
" print(f\" Pct > 2σ: {(abs(z_score) > 2).mean()*100:.1f}%\")\n",
" print(f\" Pct > 1.5σ: {(abs(z_score) > 1.5).mean()*100:.1f}%\")\n",
" \n",
" # Half-life of mean reversion\n",
" spread = ratio.dropna()\n",
" spread_lag = spread.shift(1).dropna()\n",
" spread_diff = spread - spread_lag\n",
" spread_diff = spread_diff.iloc[1:]\n",
" spread_lag = spread_lag.iloc[:len(spread_diff)]\n",
" if len(spread_lag) > 0:\n",
" import statsmodels.api as sm # may need install\n",
" try:\n",
" X = sm.add_constant(spread_lag.values)\n",
" model = sm.OLS(spread_diff.values, X).fit()\n",
" hl = -np.log(2) / model.params[1] if model.params[1] < 0 else float('inf')\n",
" print(f\"\\nMean reversion half-life: {hl:.1f} bars ({hl * pd.Timedelta(hours=1).total_seconds()/3600:.1f} hours)\")\n",
" except Exception:\n",
" print(\"\\n(Install statsmodels for half-life estimation: pip install statsmodels)\")\n"
]
},
{
"cell_type": "markdown",
"id": "e92506bb",
"metadata": {},
"source": [
"## 6. Strategy Development Checklist\n",
"\n",
"Before deploying any strategy to live/papers:\n",
"\n",
"- [ ] VectorBT backtest on real data (not synthetic)\n",
"- [ ] At least 50 trades in the backtest\n",
"- [ ] Walk-forward consistency > 50%\n",
"- [ ] DSR > 0.80, PSR > 0.70\n",
"- [ ] Haircut Sharpe > 0.50\n",
"- [ ] Maximum drawdown < 15%\n",
"- [ ] Win rate > 50% OR profit factor > 1.5\n",
"- [ ] Average trade PnL > 2x fee cost\n",
"- [ ] Correlation < 0.7 with existing portfolio strategies\n",
"- [ ] Phase 3 queue simulation (queue-aware fills) for maker strategies\n",
"- [ ] Paper trading for at least 24h before live\n",
"\n",
"**Only deploy strategies that pass all 11 checks.**"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"name": "python",
"version": "3.13.0"
}
},
"nbformat": 4,
"nbformat_minor": 5
}