Major refactor: Fantabeto 26/27 — modular package, GBM ensemble, MILP/MCTS optimization

Phase 1: Data Engineering
- Refactored notebooks into src/{scraper,features,models,optimization,bot}
- FBref scraper with proxy rotation + Playwright Cloudflare bypass
- Fantacalcio.it integrated scraper (authenticated API + HTML fallback)
- api-football RapidAPI client for supplementary xG/xA/injuries
- RAG news pipeline: Gazzetta, Sky Sport, Di Marzio → injury/suspension/tactical extraction
- 26/27 season config: teams, scoring rules, name mappings, news sources

Phase 2: SOTA ML Architecture
- GBM Ensemble (LightGBM + CatBoost + XGBoost) with stacked blending
- Bootstrap ensemble for uncertainty quantification
- SinhArcsinh distribution head (ported from original TF Probability)
- Card classifiers (yellow/red), penalty model, goal probability (Poisson)
- Temporal GNN for player interaction modeling (crosses→goals, passes→assists)
- Optuna hyperparameter tuning with time-series CV

Phase 3: Operations Research
- Auction solver: MILP knapsack with PuLP (budget + role constraints)
- Grid Auction (Asta a Griglia): Minimax game theory bidding strategy
- Weekly lineup optimizer: MCTS maximizing win probability vs opponent
- Modificatore Difesa integration + captain selection
- Transfer market analyzer: buy-low/sell-high via xG regression to mean
- Opponent behavior modeling from historical lineage patterns

Phase 4: Agentic Workflow
- Telegram bot: auto-briefing (Friday + Sunday morning)
- Tactical briefing generator with start/sit recommendations
- GitHub Actions CI/CD: scheduled pipeline (scrape → predict → notify)

Infrastructure:
- 31 pytest unit tests (features, models, scraper, optimization)
- requirements.txt (lightgbm, catboost, xgboost, optuna, pulp, playwright, langchain)
- Makefile with install/test/lint/scrape/train/bot targets
- Jupyter notebook: 26_27_strategy.ipynb demonstrating auction + matchday 1 mockup
- Completely rewritten README.md with architecture diagram
This commit is contained in:
ramseshk
2026-08-11 13:16:07 +08:00
parent 6d9167596c
commit 3b065775f5
44 changed files with 5754 additions and 57 deletions
+118
View File
@@ -0,0 +1,118 @@
"""api-football integration via RapidAPI.
Supplements FBref data with real-time injury info, expected goals (xG),
expected assists (xA), and fixture data.
Requires RAPIDAPI_KEY environment variable.
"""
import logging
import os
import time
from pathlib import Path
from typing import Optional
import pandas as pd
import requests
logger = logging.getLogger(__name__)
API_BASE = "https://api-football-v1.p.rapidapi.com/v3"
LEAGUE_ID = 135 # Serie A
class APIFootballClient:
def __init__(self, api_key: Optional[str] = None):
self.api_key = api_key or os.getenv("RAPIDAPI_KEY")
if not self.api_key:
logger.warning("No RAPIDAPI_KEY found. API calls will fail.")
self.session = requests.Session()
self.session.headers.update({
"x-rapidapi-key": self.api_key or "",
"x-rapidapi-host": "api-football-v1.p.rapidapi.com",
})
def _get(self, endpoint: str, params: Optional[dict] = None) -> dict:
if not self.api_key:
raise ValueError("RAPIDAPI_KEY not configured")
url = f"{API_BASE}/{endpoint}"
resp = self.session.get(url, params=params, timeout=15)
resp.raise_for_status()
data = resp.json()
if data.get("errors"):
logger.error(f"API error: {data['errors']}")
return data
def get_fixtures(self, season: int, matchday: Optional[int] = None) -> pd.DataFrame:
"""Get Serie A fixtures for a season. Optionally filter by matchday."""
params = {"league": LEAGUE_ID, "season": season}
if matchday:
params["round"] = f"Regular Season - {matchday}"
data = self._get("fixtures", params)
fixtures = data.get("response", [])
rows = []
for fix in fixtures:
f = fix["fixture"]
teams = fix["teams"]
rows.append({
"fixture_id": f["id"],
"date": f["date"],
"matchday": f.get("round", "").replace("Regular Season - ", ""),
"home_team": teams["home"]["name"],
"away_team": teams["away"]["name"],
"home_goals": fix.get("goals", {}).get("home"),
"away_goals": fix.get("goals", {}).get("away"),
})
return pd.DataFrame(rows)
def get_team_statistics(self, season: int, team_id: int) -> dict:
"""Get team-level statistics including xG, formations, etc."""
data = self._get("teams/statistics", {
"league": LEAGUE_ID, "season": season, "team": team_id,
})
return data.get("response", {})
def get_player_statistics(self, season: int, team_id: int, page: int = 1) -> pd.DataFrame:
"""Get player statistics for a team in a given season."""
data = self._get("players", {
"league": LEAGUE_ID, "season": season, "team": team_id, "page": page,
})
players = data.get("response", [])
rows = []
for p in players:
player = p["player"]
stats = p["statistics"][0] if p.get("statistics") else {}
rows.append({
"player_id": player["id"],
"player_name": player["name"],
"position": stats.get("games", {}).get("position", ""),
"appearences": stats.get("games", {}).get("appearences", 0),
"minutes": stats.get("games", {}).get("minutes", 0),
"goals": stats.get("goals", {}).get("total", 0),
"assists": stats.get("goals", {}).get("assists", 0),
"yellow_cards": stats.get("cards", {}).get("yellow", 0),
"red_cards": stats.get("cards", {}).get("red", 0),
})
return pd.DataFrame(rows)
def get_injuries(self, season: int, team_id: Optional[int] = None) -> pd.DataFrame:
"""Get current injury list."""
params = {"league": LEAGUE_ID, "season": season}
if team_id:
params["team"] = team_id
data = self._get("injuries", params)
injuries = data.get("response", [])
rows = []
for inj in injuries:
player = inj["player"]
row = {
"player_id": player["id"],
"player_name": player["name"],
"team": inj["team"]["name"],
"injury_type": inj.get("player", {}).get("type", ""),
"reason": inj.get("player", {}).get("reason", ""),
}
rows.append(row)
return pd.DataFrame(rows)