Major refactor: Fantabeto 26/27 — modular package, GBM ensemble, MILP/MCTS optimization
Phase 1: Data Engineering
- Refactored notebooks into src/{scraper,features,models,optimization,bot}
- FBref scraper with proxy rotation + Playwright Cloudflare bypass
- Fantacalcio.it integrated scraper (authenticated API + HTML fallback)
- api-football RapidAPI client for supplementary xG/xA/injuries
- RAG news pipeline: Gazzetta, Sky Sport, Di Marzio → injury/suspension/tactical extraction
- 26/27 season config: teams, scoring rules, name mappings, news sources
Phase 2: SOTA ML Architecture
- GBM Ensemble (LightGBM + CatBoost + XGBoost) with stacked blending
- Bootstrap ensemble for uncertainty quantification
- SinhArcsinh distribution head (ported from original TF Probability)
- Card classifiers (yellow/red), penalty model, goal probability (Poisson)
- Temporal GNN for player interaction modeling (crosses→goals, passes→assists)
- Optuna hyperparameter tuning with time-series CV
Phase 3: Operations Research
- Auction solver: MILP knapsack with PuLP (budget + role constraints)
- Grid Auction (Asta a Griglia): Minimax game theory bidding strategy
- Weekly lineup optimizer: MCTS maximizing win probability vs opponent
- Modificatore Difesa integration + captain selection
- Transfer market analyzer: buy-low/sell-high via xG regression to mean
- Opponent behavior modeling from historical lineage patterns
Phase 4: Agentic Workflow
- Telegram bot: auto-briefing (Friday + Sunday morning)
- Tactical briefing generator with start/sit recommendations
- GitHub Actions CI/CD: scheduled pipeline (scrape → predict → notify)
Infrastructure:
- 31 pytest unit tests (features, models, scraper, optimization)
- requirements.txt (lightgbm, catboost, xgboost, optuna, pulp, playwright, langchain)
- Makefile with install/test/lint/scrape/train/bot targets
- Jupyter notebook: 26_27_strategy.ipynb demonstrating auction + matchday 1 mockup
- Completely rewritten README.md with architecture diagram
This commit is contained in:
@@ -0,0 +1,55 @@
|
||||
# Fantabeto 2026/27 — Core dependencies
|
||||
python-dateutil>=2.8
|
||||
requests>=2.31
|
||||
beautifulsoup4>=4.12
|
||||
lxml>=5.0
|
||||
|
||||
# Data processing
|
||||
pandas>=2.1
|
||||
numpy>=1.26
|
||||
openpyxl>=3.1
|
||||
scipy>=1.11
|
||||
pyyaml>=6.0
|
||||
|
||||
# ML models
|
||||
lightgbm>=4.3
|
||||
catboost>=1.2
|
||||
xgboost>=2.0
|
||||
scikit-learn>=1.4
|
||||
|
||||
# Hyperparameter tuning
|
||||
optuna>=3.5
|
||||
|
||||
# Graph neural network (optional, for T-GNN)
|
||||
torch>=2.2
|
||||
torch-geometric>=2.5
|
||||
|
||||
# Optimization
|
||||
pulp>=2.8
|
||||
|
||||
# Browser automation (Playwright fallback for Cloudflare)
|
||||
playwright>=1.42
|
||||
|
||||
# RAG & LLM (optional, for news pipeline)
|
||||
langchain>=0.1
|
||||
langchain-community>=0.1
|
||||
chromadb>=0.4
|
||||
feedparser>=6.0
|
||||
|
||||
# LLM providers (choose one)
|
||||
openai>=1.12
|
||||
# anthropic>=0.20
|
||||
# ollama>=0.1
|
||||
|
||||
# Visualization
|
||||
matplotlib>=3.8
|
||||
seaborn>=0.13
|
||||
plotly>=5.18
|
||||
|
||||
# Development
|
||||
pytest>=8.0
|
||||
black>=24.0
|
||||
ruff>=0.3
|
||||
|
||||
# Bot
|
||||
python-telegram-bot>=21.0
|
||||
Reference in New Issue
Block a user