Files
ramseshk 3b065775f5 Major refactor: Fantabeto 26/27 — modular package, GBM ensemble, MILP/MCTS optimization
Phase 1: Data Engineering
- Refactored notebooks into src/{scraper,features,models,optimization,bot}
- FBref scraper with proxy rotation + Playwright Cloudflare bypass
- Fantacalcio.it integrated scraper (authenticated API + HTML fallback)
- api-football RapidAPI client for supplementary xG/xA/injuries
- RAG news pipeline: Gazzetta, Sky Sport, Di Marzio → injury/suspension/tactical extraction
- 26/27 season config: teams, scoring rules, name mappings, news sources

Phase 2: SOTA ML Architecture
- GBM Ensemble (LightGBM + CatBoost + XGBoost) with stacked blending
- Bootstrap ensemble for uncertainty quantification
- SinhArcsinh distribution head (ported from original TF Probability)
- Card classifiers (yellow/red), penalty model, goal probability (Poisson)
- Temporal GNN for player interaction modeling (crosses→goals, passes→assists)
- Optuna hyperparameter tuning with time-series CV

Phase 3: Operations Research
- Auction solver: MILP knapsack with PuLP (budget + role constraints)
- Grid Auction (Asta a Griglia): Minimax game theory bidding strategy
- Weekly lineup optimizer: MCTS maximizing win probability vs opponent
- Modificatore Difesa integration + captain selection
- Transfer market analyzer: buy-low/sell-high via xG regression to mean
- Opponent behavior modeling from historical lineage patterns

Phase 4: Agentic Workflow
- Telegram bot: auto-briefing (Friday + Sunday morning)
- Tactical briefing generator with start/sit recommendations
- GitHub Actions CI/CD: scheduled pipeline (scrape → predict → notify)

Infrastructure:
- 31 pytest unit tests (features, models, scraper, optimization)
- requirements.txt (lightgbm, catboost, xgboost, optuna, pulp, playwright, langchain)
- Makefile with install/test/lint/scrape/train/bot targets
- Jupyter notebook: 26_27_strategy.ipynb demonstrating auction + matchday 1 mockup
- Completely rewritten README.md with architecture diagram
2026-08-11 13:16:07 +08:00

81 lines
2.4 KiB
YAML

name: Fantabeto Weekly Pipeline
on:
schedule:
- cron: '0 18 * * 5' # Friday 18:00 UTC = 20:00 CET
- cron: '0 8 * * 0' # Sunday 08:00 UTC = 10:00 CET
workflow_dispatch: # Manual trigger
jobs:
run-pipeline:
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Set up Python 3.11
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Cache pip packages
uses: actions/cache@v4
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-${{ hashFiles('requirements.txt') }}
restore-keys: |
${{ runner.os }}-pip-
- name: Install dependencies
run: |
pip install --upgrade pip
pip install -r requirements.txt
- name: Run tests
run: python -m pytest tests/ -v --tb=short
- name: Scrape latest data
run: |
python -c "
from src.scraper.fbref_scraper import scrape_current_season
scrape_current_season('data/fbref')
"
env:
FANTACALCIO_TOKEN: ${{ secrets.FANTACALCIO_TOKEN }}
RAPIDAPI_KEY: ${{ secrets.RAPIDAPI_KEY }}
- name: Build features
run: python -c "from src.pipeline import Pipeline; p = Pipeline(); p.build_features()"
- name: Run predictions
run: |
mkdir -p data/predictions
python -c "
import pandas as pd
from src.pipeline import Pipeline
p = Pipeline()
df = pd.read_excel('data/match_dataset.xlsx')
model = p.train_models(df.drop(columns=['fantavote','vote'], errors='ignore'), df['fantavote'])
"
- name: Send Telegram briefing
env:
TELEGRAM_BOT_TOKEN: ${{ secrets.TELEGRAM_BOT_TOKEN }}
TELEGRAM_CHAT_ID: ${{ secrets.TELEGRAM_CHAT_ID }}
run: |
python -c "
from src.bot.telegram_bot import TelegramBot
from src.bot.briefing import BriefingGenerator
bot = TelegramBot()
gen = BriefingGenerator()
bot.send_briefing('Fantabeto 26/27 weekly pipeline completed. Predictions ready.')
"
- name: Upload predictions artifact
uses: actions/upload-artifact@v4
with:
name: predictions
path: data/predictions/