{"id":50262963,"url":"https://github.com/yupoet/aurumq-rl","last_synced_at":"2026-06-13T03:00:42.472Z","repository":{"id":354316695,"uuid":"1222572183","full_name":"yupoet/aurumq-rl","owner":"yupoet","description":"RL stock selection for China A-share — bundled polars-native factor library (105 Alpha101 + 191 GTJA Alpha191 = 296 factors), board-aware price limits, GPU train + ONNX CPU infer, MIT-licensed.","archived":false,"fork":false,"pushed_at":"2026-05-14T03:57:02.000Z","size":11360,"stargazers_count":3,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-05-14T05:37:33.847Z","etag":null,"topics":["a-share","algorithmic-trading","alpha101","alpha191","china-stock-market","factor-investing","factor-library","gtja191","guotai-junan","gymnasium","onnx","polars","ppo","pytorch","quantitative-finance","quantitative-trading","reinforcement-learning","stable-baselines3","stock-selection"],"latest_commit_sha":null,"homepage":"https://github.com/yupoet/aurumq-rl","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/yupoet.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-04-27T13:50:53.000Z","updated_at":"2026-05-14T03:57:06.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/yupoet/aurumq-rl","commit_stats":null,"previous_names":["yupoet/aurumq-rl"],"tags_count":2,"template":false,"template_full_name":null,"purl":"pkg:github/yupoet/aurumq-rl","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yupoet%2Faurumq-rl","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yupoet%2Faurumq-rl/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yupoet%2Faurumq-rl/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yupoet%2Faurumq-rl/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/yupoet","download_url":"https://codeload.github.com/yupoet/aurumq-rl/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yupoet%2Faurumq-rl/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34270417,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-13T02:00:06.617Z","response_time":62,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["a-share","algorithmic-trading","alpha101","alpha191","china-stock-market","factor-investing","factor-library","gtja191","guotai-junan","gymnasium","onnx","polars","ppo","pytorch","quantitative-finance","quantitative-trading","reinforcement-learning","stable-baselines3","stock-selection"],"created_at":"2026-05-27T12:00:25.411Z","updated_at":"2026-06-13T03:00:42.465Z","avatar_url":"https://github.com/yupoet.png","language":"Python","funding_links":[],"categories":["Trading \u0026 Backtesting"],"sub_categories":[],"readme":"# AurumQ-RL · A股量化强化学习选股开源项目\n# AurumQ-RL · An Open-Source Reinforcement Learning Stock-Selection Framework for the China A-Share Market\n\n\u003e 中文：一份面向 A 股的因子工程 + 强化学习选股参考实现，附完整的迭代史、消融实验、生产化决策与教训。\n\u003e English: A factor-engineering + reinforcement-learning stock-picking reference implementation for China A-shares, shipped with the full iteration history, ablations, productionization decisions, and lessons learned.\n\n\u003cp align=\"left\"\u003e\n  \u003ca href=\"https://github.com/yupoet/aurumq-rl/actions/workflows/ci.yml\"\u003e\u003cimg src=\"https://github.com/yupoet/aurumq-rl/actions/workflows/ci.yml/badge.svg\" alt=\"CI\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/yupoet/aqml\"\u003e\u003cimg src=\"https://img.shields.io/badge/AQML-strategy%20DSL-D4A853.svg\" alt=\"AQML strategy DSL\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/yupoet/aurumq-rl/blob/main/LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/badge/license-MIT-blue.svg\" alt=\"License: MIT\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://www.python.org/downloads/\"\u003e\u003cimg src=\"https://img.shields.io/badge/python-3.10%20%7C%203.11%20%7C%203.12-blue.svg\" alt=\"Python 3.10+\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/yupoet/aurumq-rl/releases\"\u003e\u003cimg src=\"https://img.shields.io/github/v/release/yupoet/aurumq-rl?include_prereleases\u0026label=release\" alt=\"Latest release\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/yupoet/aurumq-rl\"\u003e\u003cimg src=\"https://img.shields.io/badge/tests-1386%20passed-brightgreen.svg\" alt=\"Tests: 1386 passed\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/yupoet/aurumq-rl/tree/main/docs/factor_library\"\u003e\u003cimg src=\"https://img.shields.io/badge/factors-105%20alpha101%20%2B%20191%20gtja191-purple.svg\" alt=\"Factor library\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://stable-baselines3.readthedocs.io/\"\u003e\u003cimg src=\"https://img.shields.io/badge/RL-Stable--Baselines3-blueviolet.svg\" alt=\"Stable-Baselines3\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://onnxruntime.ai/\"\u003e\u003cimg src=\"https://img.shields.io/badge/inference-ONNX%20Runtime-yellow.svg\" alt=\"ONNX Runtime\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/astral-sh/ruff\"\u003e\u003cimg src=\"https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json\" alt=\"Ruff\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/yupoet/aurumq-rl/stargazers\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/yupoet/aurumq-rl?style=social\" alt=\"GitHub stars\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/assets/social-preview.png\" alt=\"AurumQ-RL — Reinforcement learning stock selection for China A-share market\" width=\"720\"\u003e\n\u003c/p\u003e\n\n\u003cp align=\"left\"\u003e\n  \u003cstrong\u003e📊 China A-share · 🤖 PPO/A2C/SAC · 🚀 GPU Train + CPU Infer · 📈 Alpha101 + GTJA Alpha191 + Main-Force + Hot-Money + Northbound · 🧪 26 Phases of Open Experiments\u003c/strong\u003e\n\u003c/p\u003e\n\n---\n\n## 摘要 / Abstract\n\n**中文.** AurumQ-RL 是一个针对 A 股市场特有微观结构（T+1、±10% 涨跌停、主板/科创/创业/北交分层、ST 风险警示、申万一级行业、龙虎榜、北向、游资席位、筹码分布）做工程化封装的强化学习选股开源项目。仓库内含：(1) 一个 polars-native 的因子计算引擎，覆盖 105 个 WorldQuant Alpha101 + 191 个国泰君安 Alpha191（合计 296 个量价因子）外加 11 个 A 股私有因子族（mf_, mfp_, hm_, hk_, inst_, mg_, cyq_, senti_, sh_, fund_, ind_, mkt_, gtja_, tech_, cmf_, zt_）；(2) Stable-Baselines3 PPO 训练栈，针对 RTX 4070 12 GB 做了 GPU 化重构（per-stock 编码器 + CUDA-resident rollout buffer + 索引化观测）；(3) 完整的 14-phase 训练栈演化史，从最初 11% GPU 利用率到 1M-step 隔夜训练；(4) 26-phase 模型实验史，覆盖奖励重设计、长 panel 消融、rank-z 假设检验、SHAP 剪枝、事件衰减编码等关键转折；(5) ONNX 导出 + CPU 推理生产管线。**核心实证发现**：(a) rank-z 跨截面归一化会在长 panel 训练中销毁 5-6 bps 的跨年因子幅度信号；(b) 5 年训练窗口已是 plateau，2018-2019 数据零边际贡献；(c) Strategy D（top-K 仓位按分数加权）能与任何基模型叠加 +7-10 bps 的 mean_y；(d) 二值事件标志直接进 LayerNorm 是 −33% 准确率回归的元凶，必须用 exp-decay τ=10d 编码；(e) cyq 筹码因子的回填 vs 实采分布漂移导致 1.5× T-1 hit 回归，根因是 z-score 不抹平时序 regime 跳变。**当前生产状态（2026-05-11）**：10 个 model_version 同时在 Celery Beat 18:50-19:00 排程，新进 best 是 path5_long（H1 校准 mean_y +0.02882，T1_hit 55.8%）。\n\n**English.** AurumQ-RL is an open-source RL stock-selection framework engineered around the China A-share market's unique microstructure (T+1 settlement; ±10% daily limit per board; main-board vs ChiNext/STAR/BSE segmentation; ST risk-warning; SW Tier-1 industry; Dragon-Tiger List; Northbound; hot-money seats; chip distribution). The repository ships (1) a polars-native factor engine covering 105 WorldQuant Alpha101 + 191 Guotai Junan Alpha191 (296 price-volume factors total) plus 11 private A-share factor families (mf_, mfp_, hm_, hk_, inst_, mg_, cyq_, senti_, sh_, fund_, ind_, mkt_, gtja_, tech_, cmf_, zt_); (2) a Stable-Baselines3 PPO training stack with a GPU-vectorized rewrite for the RTX 4070 12 GB (per-stock encoder + CUDA-resident rollout buffer + index-only observations); (3) the complete 14-phase training-stack evolution from 11 % GPU utilization at bring-up to overnight 1M-step training; (4) the 26-phase modeling experiment history covering reward redesign, long-panel ablations, the rank-z hypothesis test, SHAP-based pruning, and event-decay encoding; (5) ONNX export + CPU inference production pipeline. **Key empirical findings.** (a) Per-day cross-sectional rank-z destroys 5-6 bps of cross-year factor amplitude signal in long-panel training. (b) Five-year training window is the plateau; 2018-2019 data contribute nothing. (c) Strategy D top-K score-weighted sizing compounds +7-10 bps mean_y onto any base model. (d) Binary event flags fed directly into LayerNorm cause a −33 % accuracy regression and must be encoded as exp-decay with τ=10d. (e) Chip-distribution (cyq) backfill versus real-sample distribution shift drove a 1.5× T-1 hit regression — root cause is that cross-section z-score does not equalize mid-stream temporal regime shifts. **Current production state (2026-05-11):** 10 model versions co-scheduled in the Celery Beat 18:50-19:00 budget; the new best is `path5_long` (H1 calibrated mean_y +0.02882, T1_hit 55.8 %).\n\n---\n\n## 目录 / Table of Contents\n\n[中文](#中文导览) · [English](#english-toc) · [Phase Timeline](#phase-timeline)\n\n### 中文导览\n\n- [§1 引言：为什么 A 股市场需要专门的 RL 框架](#1-引言--为什么-a-股市场需要专门的-rl-框架)\n- [§2 系统总览：数据契约、因子前缀、宇宙过滤](#2-系统总览)\n- [§3 因子库：296 量价因子 + 13 个 A 股私有因子族](#3-因子库)\n- [§4 训练方法演进：14 个 Phase 的工程史](#4-训练栈演进史--phase-0--14)\n- [§5 模型实验史：26 个 Phase 的研究决策](#5-模型实验史--phase-15--26)\n- [§6 监督学习对照赛道（paris 侧 P0/P2/Path 1-6）](#6-监督学习对照赛道paris-侧)\n- [§7 实证发现：6 条改变方向的结论](#7-实证发现--六条改变方向的结论)\n- [§8 生产流水线：每日 18:30-19:00 评分预算](#8-生产流水线--每日-1830-1900-评分预算)\n- [§9 工程教训：从踩坑到守则](#9-工程教训--从踩坑到守则)\n- [§10 上手与复现](#10-上手与复现)\n- [§11 路线图、引用、许可](#11-路线图引用许可)\n- [§12 研究范式分类与未来方向](#12-research-paradigms)\n\n### English TOC\n\n- [§1 Introduction: Why A-shares Need a Dedicated RL Framework](#1-introduction)\n- [§2 System Overview: Data Contract, Factor Prefixes, Universe Filter](#2-system-overview)\n- [§3 Factor Library: 296 Price-Volume + 13 Private A-share Families](#3-factor-library)\n- [§4 Training-Stack Evolution: 14 Phases of Engineering](#4-training-stack-evolution--phase-0-to-14)\n- [§5 Modeling Experiment History: 26 Phases of Research Decisions](#5-modeling-experiment-history--phase-15-to-26)\n- [§6 Supervised-Learning Companion Track (paris side P0/P2/Path 1-6)](#6-supervised-learning-companion-track)\n- [§7 Empirical Findings: Six Conclusions That Changed Direction](#7-empirical-findings)\n- [§8 Production Pipeline: Daily 18:30-19:00 Scoring Budget](#8-production-pipeline)\n- [§9 Engineering Lessons: From Pitfalls to Operating Rules](#9-engineering-lessons)\n- [§10 Quick Start and Reproduction](#10-quick-start-and-reproduction)\n- [§11 Roadmap, Citation, License](#11-roadmap-citation-license)\n- [§12 Research Paradigms and Future Directions](#12-research-paradigms)\n\n### Phase Timeline\n\n```\nPhase 0       Synthetic pipeline-up                 (~pre-2026-04-29)\nPhase 1       First real-data alpha101 PPO          2026-04-29..30\nPhase 2       Combined 355-col panel + wider net    2026-04-30\nPhase 3       R1/R2/R3 smoke-round tuning           2026-05-01 morning\nPhase 4       fps scaling, IPC ceiling discovery    2026-05-01 noon\nPhase 5       Realizations → GPU framework redesign 2026-05-01 PM\nPhase 6/7     GPU-vectorized framework + 50k smoke  2026-05-01 evening\nPhase 8       GPURolloutBuffer (CUDA-resident)      2026-05-01 evening\nPhase 9       IndexOnlyRolloutBuffer + n_steps=1024 2026-05-01 late evening\nPhase 10      Optimizer orphan + LayerNorm + dual pooling  2026-05-01 night\nPhase 11/12   bf16 autocast / target_kl=0.10 (eliminated)\nPhase 13      PPO SGD perf-probe                    2026-05-01 late night\nPhase 14      TF32 + unique-date + 1M overnight     2026-05-02 → 03 early\nPhase 15      RL serving integration (champion)     2026-05-02..03\nPhase 16-19   Eval correction + multi-seed ensemble 2026-05-03\nPhase 20      Long-data PPO                         2026-05-05\nPhase 21      V2 forward_10d REJECTED               2026-05-05..06\nPhase 22      Main-wave reward redesign             2026-05-06\nPhase 23      Episode targets cleanup               2026-05-06\nPhase 24/25   Tech-factor + importance-weight REJECT 2026-05-07\nPhase 26A-G   cyq fix + event-decay tech (26F prod) 2026-05-07..08\n\n# SL companion (paris side, AurumQ repo)\nP0            Wave label ablation (Method A wins)   2026-05-09\nP2            5-seed ensemble (wave_t3_lgbm_v2)     2026-05-09\nP3            PPO residual ALGORITHM_SPEC v2        2026-05-09\nPath 1-6      Multi-path SL exploration              2026-05-10\nLong-panel    Hybrid / path1_long / path5_long      2026-05-10..11\n```\n\n---\n\n## §1 引言 — 为什么 A 股市场需要专门的 RL 框架\n## §1 Introduction — Why A-shares Need a Dedicated RL Framework \u003ca id=\"1-introduction\"\u003e\u003c/a\u003e\n\n### 1.1 市场微观结构 / Market Microstructure\n\n**中文.** A 股市场和欧美/港股/加密的微观结构差异远大于表面上的\"也是股票\"。本项目所有训练、回测、因子计算一律仅在 A 股**主板 + 非 ST + 未退市**（约 3000 只）上跑，这是 CLAUDE.md 的硬约束。理由是 regime homogeneity：\n\n- **板别价格限制差异**：主板 ±10%、ST/*ST ±5%、科创板 ±20%、创业板 ±20%、北交所 ±30%。把这五个池子混进同一个训练 batch，模型必须额外学一个\"我现在在哪个板\"的元变量，跨年泛化崩溃。\n- **T+1 结算**：当日买入次日才能卖。RL 训练时的\"动作-奖励\"延迟和欧美/加密 T+0 完全不同；用 t→t+1 的同日 PnL 等价于看穿未来。\n- **集合竞价 vs 连续竞价**：9:15-9:30 集合竞价是 A 股最重要的\"标价\"事件之一，`stk_auction.bid_*` 字段在 `kpl_list` 里需要单独抓，常规 `pct_chg/amount` 不适用。\n- **披露节奏**：主板财报 + 业绩预告时点高度同步，跨年信号同质性强；科创板/北交所披露规则差异显著。\n- **流动性结构**：主板日均成交额 vs 北交所差 1-2 个数量级，分位归一化（rank-z）会把北交所的微弱信号放大到和主板同尺度，污染监督信号。\n\n**English.** A-share microstructure differs from US/HK/crypto markets in ways that defeat naive \"stocks are stocks\" transfer. All training, backtest, and factor computation are restricted by hard constraint (CLAUDE.md) to **main-board, non-ST, non-delisted** stocks (~3000 names). Reasons:\n\n- **Board-specific price limits**: main-board ±10 %, ST/*ST ±5 %, STAR/ChiNext ±20 %, BSE ±30 %. Mixing these into a single training batch forces the model to learn a \"which board am I on\" meta-variable, destroying cross-year generalization.\n- **T+1 settlement**: bought today cannot be sold until tomorrow. The action-reward lag in RL training is fundamentally different from T+0 US/crypto. Using same-day t→t+1 PnL is look-ahead.\n- **Call auction vs continuous trading**: the 09:15–09:30 call auction is one of the most informative tape events in A-shares. `stk_auction.bid_*` fields in `kpl_list` need special extraction; standard `pct_chg/amount` do not apply.\n- **Disclosure cadence**: main-board earnings + pre-announcement timing is tightly synchronized, giving cross-year signal homogeneity; STAR/BSE differ.\n- **Liquidity structure**: main-board daily turnover vs BSE differs by 1–2 orders of magnitude; rank-z normalization would inflate BSE micro-signals to the same scale as main-board, polluting supervision.\n\n### 1.2 为什么 RL 而不是因子排序 / Why RL Rather Than Linear Alpha Aggregation\n\n**中文.** 传统量化做法把 alpha101/gtja191 当 96 个排名分数线性加权，找最优权重向量。问题：\n\n1. **截面 IC 是低维度量**：300 个因子 × 200 天 = 6 万样本估 300 个权重，过拟合容易，跨年泛化差。\n2. **非线性交互被忽略**：`alpha_026 × cyq_winning_ratio` 在 30% 行业暴露上限下的表现可能远超两者线性和，线性模型抓不到。\n3. **奖励/成本不进训练目标**：传统模型最小化 IC residual，但真实回报是「扣除 30bp 双边费 + T+1 不能反手 + 单行业 30% 上限」之后的实现收益，目标函数和评估指标不一致。\n4. **状态依赖动作**：今天选哪 50 只取决于全市场截面分布，不是任意 200 只独立打分加总。\n\nRL 用神经网络直接学映射「当前因子截面 + 持仓状态 → top-K 动作」，把成本、流动性、约束都搬进环境。代价是：样本效率低、调参成本高、可解释性差。本项目实事求是地认为「**RL 不一定比线性 alpha 加权强**」，因此并行维护一条监督学习（SL）赛道作为对照（见 §6）；事实证明 SL 赛道的 `path5_long` 当前是综合最佳（H1 +0.02882），但 RL 赛道的 `phase15e_150k_grand_champion` 在 Sharpe 维度（OOS +6.27）仍是不可替代的多样性来源。\n\n**English.** Conventional quant treats alpha101/gtja191 as ~96 ranking scores to linearly combine. Problems:\n\n1. **Cross-sectional IC is a low-dimensional statistic.** 300 factors × 200 days = 60k samples to estimate 300 weights — overfitting prone, poor cross-year generalization.\n2. **Nonlinear interactions go unmodeled.** `alpha_026 × cyq_winning_ratio` under a 30 %-industry-cap can dominate either factor alone; linear models cannot capture this.\n3. **Cost/constraint are absent from the training objective.** Traditional models minimize IC residual; realized return is net of 30bp round-trip cost, T+1 inability to reverse, and 30 % single-industry exposure cap. Objective and evaluation diverge.\n4. **State-dependent action.** Today's top-50 picks depend on the full market cross-section, not 200 independent scores summed.\n\nRL learns the mapping `current factor cross-section + holdings → top-K action` directly, moving cost / liquidity / constraints into the environment. Tradeoffs: poor sample efficiency, expensive hyperparameter tuning, weak interpretability. This project honestly does NOT assume **RL is always better than linear alpha aggregation** — we maintain a parallel supervised-learning (SL) track as control (§6). Today the SL track's `path5_long` is the all-around best (H1 +0.02882), but the RL track's `phase15e_150k_grand_champion` remains an irreplaceable Sharpe-dimension diversity contributor (OOS +6.27).\n\n### 1.3 项目定位 / Scope and Boundaries\n\n**中文.** 本项目**是**：\n- ✅ RL 选股算法 + Gymnasium 环境的参考实现\n- ✅ A 股微观结构（T+1 / 涨跌停 / ST / 板别 / 行业暴露上限）的工程化封装\n- ✅ 多源因子的**消费者**（按列名前缀识别，输入有什么就用什么）\n- ✅ 离线训练 → ONNX 导出 → CPU 推理的端到端流水线\n- ✅ **完整的研究决策记录**：每个 phase 都记录了 stack 变更 / 量化证据 / 拒绝原因，便于复现和审计\n\n**不是**：\n- ❌ 实盘交易系统（无券商接口、无下单 API）\n- ❌ 数据采集工具（不内置任何数据 API key；用户自己 pipeline 写 Parquet）\n- ❌ 因子计算库（因子计算可选——`aurumq_rl.factors` 提供 296 个 polars 实现，但用户可以完全自己算）\n- ❌ 高频交易（日频选股，T+1 持仓）\n\n**English.** This project **is**:\n- ✅ Reference implementation of an RL stock-picking algorithm + Gymnasium environment\n- ✅ Engineering encapsulation of A-share microstructure (T+1 / price limits / ST / board / industry cap)\n- ✅ A **consumer** of multi-source factors (prefix-recognized; uses whatever is in your Parquet)\n- ✅ End-to-end offline-train → ONNX-export → CPU-infer pipeline\n- ✅ **Complete research decision log**: every phase records stack diff, quantitative evidence, rejection reason — auditable and reproducible\n\n**Is NOT**:\n- ❌ A live trading system (no broker adapter, no order API)\n- ❌ A data-ingestion tool (no API keys bundled; you write the Parquet)\n- ❌ A factor-computation library (optional — `aurumq_rl.factors` offers 296 polars implementations, but you can roll your own)\n- ❌ High-frequency trading (daily-frequency stock picking under T+1)\n\n### 1.4 AurumQ 生态 / The AurumQ Ecosystem\n\n**中文.** 本项目是 AurumQ 平台的一个开源子模块，整个生态分三层：\n\n| 层 | 项目 | 职责 |\n|---|---|---|\n| 策略 DSL | [AQML](https://github.com/yupoet/aqml) | `.aqml` 声明式策略，可读 / 可验证 / AI 可生成的筛选 + 打分 + 风控规则 |\n| 因子 + RL | **AurumQ-RL** (本项目) | 因子工程、A 股约束、PPO/A2C/SAC、ONNX 推理 |\n| 平台 | AurumQ (闭源) | Web 平台 + REST API + 模拟盘 + 风控引擎 + AI 投研 |\n\n典型工作流：先用 AQML 写策略意图 → 用 AurumQ-RL 把因子列和约束塞进模型训练 → 训出的 `.onnx` 回到 AurumQ 平台跑模拟盘和实时排程。\n\n**English.** This project is an open-source submodule of the AurumQ platform; the ecosystem has three layers:\n\n| Layer | Project | Role |\n|---|---|---|\n| Strategy DSL | [AQML](https://github.com/yupoet/aqml) | `.aqml` declarative strategy — human-readable, validatable, AI-generatable screening + scoring + risk rules |\n| Factors + RL | **AurumQ-RL** (this repo) | Factor engineering, A-share constraints, PPO/A2C/SAC, ONNX inference |\n| Platform | AurumQ (proprietary) | Web platform + REST API + paper trading + risk engine + AI research |\n\nTypical workflow: write strategy intent in AQML → feed factor columns and constraints into AurumQ-RL for training → the resulting `.onnx` returns to the AurumQ platform for paper trading and real-time scheduling.\n\n---\n\n## §2 系统总览 \u003ca id=\"2-系统总览\"\u003e\u003c/a\u003e\n## §2 System Overview \u003ca id=\"2-system-overview\"\u003e\u003c/a\u003e\n\n### 2.1 数据契约 / The Data Contract\n\n**中文.** 项目对外契约就一句话：**给我一份 Parquet，我就能训练**。Parquet 必含：\n\n- `ts_code` (str): Tushare 风格代码 `XXXXXX.SH/SZ/BJ`\n- `trade_date` (date): 交易日\n- `close` (float): 收盘价\n- `pct_chg` (float): 涨跌幅（**小数形式**，+10% = 0.10，不是 10.0）\n- `vol` (float): 成交量（== 0 视为停牌）\n- 因子列（至少包含一组前缀）：`alpha_*` / `gtja_*` / `mf_*` / `mfp_*` / `hm_*` / `hk_*` / `inst_*` / `mg_*` / `cyq_*` / `senti_*` / `sh_*` / `fund_*` / `ind_*` / `mkt_*` / `tech_*` / `cmf_*` / `zt_*`\n\n**可选字段**（提供则使用，不提供则自动降级）：\n\n- `is_st` (bool): ST 标记，缺则按全 False 处理\n- `days_since_ipo` (int): 上市以来交易日数（用于新股 60 日保护）\n- `industry_code` (int): 申万一级行业编码（用于 30% 行业暴露上限）\n- `is_hs300` / `is_zz500` (bool): 是否成分股，**按 trade_date 历史变更**（支持「2024-01 在 300、2024-06 调出」的时变性）\n\n数据怎么来不是本项目关心的事，三种取数方式：\n\n1. 用 `scripts/generate_synthetic.py` 一键生成 10 MB 合成数据 demo\n2. 用 `scripts/export_factor_panel.py` 从 PostgreSQL 自己的数据仓库抽取（含 SQL 模板，含 HS300/ZZ500 成员标志支持）\n3. 自己用任何工具（pandas / DuckDB / Spark）造一份满足契约的 Parquet 写入\n\n**English.** The single contract is: **give me a Parquet, I will train**. Required columns:\n\n- `ts_code` (str): Tushare-style code `XXXXXX.SH/SZ/BJ`\n- `trade_date` (date)\n- `close` (float)\n- `pct_chg` (float): **decimal form**, +10 % = 0.10, NOT 10.0\n- `vol` (float): 0 means suspended\n- Factor columns under at least one prefix: `alpha_*` / `gtja_*` / `mf_*` / `mfp_*` / `hm_*` / `hk_*` / `inst_*` / `mg_*` / `cyq_*` / `senti_*` / `sh_*` / `fund_*` / `ind_*` / `mkt_*` / `tech_*` / `cmf_*` / `zt_*`\n\n**Optional fields** (auto-fallback when absent): `is_st`, `days_since_ipo`, `industry_code`, `is_hs300`, `is_zz500` (time-varying by `trade_date`).\n\nThree ways to get data: (1) synthetic demo (`generate_synthetic.py`), (2) export from your own PG warehouse (`export_factor_panel.py`), (3) BYO Parquet from any tool that meets the contract.\n\n### 2.2 因子前缀识别 / Factor-Prefix Auto-Discovery\n\n**中文.** `data_loader.py` 通过列名前缀识别因子组，**输入 Parquet 中存在的前缀就被自动加载，不存在的自动跳过**。这套设计的核心是：项目本身不知道你给的是哪些因子，只要列名前缀对得上就一律纳入观测。\n\n| 前缀 | 含义 | 推荐维度 | 输入数据要求 |\n|---|---|---|---|\n| `alpha_*` | WorldQuant Alpha101（项目自带 105 个实现，含 6 个自定义补充）| 105 | 日频 OHLCV + amount |\n| `gtja_*` | 国泰君安 Alpha191（项目自带 191 个实现）| 191 | 日频 OHLCV + vwap + amount + 基准指数 OHLC |\n| `mf_*` | Money Flow Velocity — 主力资金流速（4 档累计筹码）| 14 + 6 `_log` 变体 | 4 档资金流分档 |\n| `mfp_*` | Main Force Position — 主力筹码持仓（与 `mf_` 互补，**不要混用**）| 12 | 主力净持仓时序表 |\n| `hm_*` | Hot Money — 主流游资席位 | 6 | 龙虎榜游资席位日成交明细 |\n| `hk_*` | Northbound — 北向资金真实持股 | 4 | 北向持股日表（港股通名单内） |\n| `inst_*` | Institutional — 龙虎榜机构净买入 | 3 | 龙虎榜机构席位明细 |\n| `mg_*` | Margin — 融资融券 | 3 | 融资融券日表 |\n| `cyq_*` | Chip Distribution — 筹码分布 | 3 | Tushare cyq_perf 表 |\n| `senti_*` | Sentiment — 涨停板情绪 | 3 | 涨停板池 + 热度榜 |\n| `sh_*` | Shareholder — 股东户数 + 大股东增减持 | 2 | 股东数据 |\n| `fund_*` | Fundamentals — 基本面 PE/PB/ROE/营收增速 | 4 | 基本面表 |\n| `ind_*` | Industry — 申万行业相对强度 | 2 | 行业指数 |\n| `mkt_*` | Market — 大盘 + 拥挤度 | 2 | 指数日表 |\n| `tech_*` | Technical — 上游算好的 MA/KDJ/MACD/Bollinger（v1.1 后 30 列）| 30 | 已 z-score 的 OHLCV 技术指标 |\n| `cmf_*` | Chaikin Money Flow — 60d/120d 累计资金流 | 2 | 量价资金流派生 |\n| `zt_*` | 涨停板 stats — 30d/60d 涨停频次、首板、连板 | 6 | 涨停板池 + 历史 |\n\n**总维度灵活**：纯 Alpha101 = 105 维 / Alpha101 + GTJA191 = 296 维 / 全部 17 前缀 ≈ 360 维 / 自定义任意子集。\n\n`StockPickingConfig.n_factors` 决定取前 N 个因子（按字母序），多余的丢弃，不足的报错。\n\n**English.** `data_loader.py` recognizes factor groups by column-name prefix; **whatever prefixes are present in your Parquet get loaded, whatever is absent gets skipped**. The design assumes the project does not know which factors you have — match the prefix, get included as observation.\n\n(See the Chinese table above for the 17-prefix breakdown. Total flexible: pure alpha101 = 105 dims / alpha+gtja = 296 / all 17 prefixes ≈ 360.)\n\n### 2.3 宇宙过滤 / Universe Filtering\n\n**中文.** 默认 `UniverseFilter.MAIN_BOARD_NON_ST` 应用六道 AND 闸门：\n\n1. **`data_ok`**: 当日有日线数据 AND `vol \u003e 0`（剔除停牌）\n2. **`main_board`**: `60[0135]\\d{3}.SH` ∪ `00[0123]\\d{3}.SZ`（剔除 300***/688***/689***/4xx/8xx/9xx）\n3. **`listed`**: `days_since_ipo ≥ 60`（新股 60 日保护）\n4. **`not_delisted`**: `stock_info.delist_date IS NULL`\n5. **`not_st`**: `is_st == False` AND stock_name 不含 \"ST\" / \"*ST\" / \"退\"\n6. **`not_suspended`**: 当日非停牌（vol \u003e 0 也涵盖此条件）\n\n应用顺序很重要：先 `data_ok` 再 `main_board`，避免在停牌日按板别 regex 算返回值时遇到 NaN/Null 行的 regex 失败。\n\n如要自定义：\n\n```python\nfrom aurumq_rl.data_loader import UniverseFilter, load_panel\n\n# 全市场（仅排 ST + 停牌）\npanel = load_panel(\"data.parquet\", universe_filter=UniverseFilter.ALL_NON_ST)\n\n# 只跑沪深 300\npanel = load_panel(\"data.parquet\", universe_filter=UniverseFilter.HS300)\n```\n\n**English.** Default `UniverseFilter.MAIN_BOARD_NON_ST` applies six AND-gates: `data_ok` (has bar + vol \u003e 0), `main_board` (regex `60[0135]\\d{3}.SH ∪ 00[0123]\\d{3}.SZ`), `listed` (days_since_ipo ≥ 60), `not_delisted`, `not_st` (`is_st = False` AND `stock_name` excludes ST/*ST/退), `not_suspended`. Ordering matters: `data_ok` before `main_board` to avoid regex on NaN. Alternative filters: `ALL_NON_ST`, `HS300`, `ZZ500`, or supply your own callable.\n\n### 2.4 模块架构 / Module Architecture\n\n**中文.** 仓库布局：\n\n```text\naurumq-rl/\n├── src/aurumq_rl/\n│   ├── env.py                  # StockPickingEnv (Gymnasium)\n│   ├── gpu_env.py              # GPU-vectorized env (Phase 6+)\n│   ├── portfolio_weight_env.py # 连续权重组合环境（马科维茨扩展）\n│   ├── data_loader.py          # Parquet → numpy/cuda 面板（多前缀识别）\n│   ├── policies/\n│   │   ├── per_stock_encoder.py    # Deep Sets 风格 per-stock 编码器\n│   │   └── shared_policy.py        # 早期 flat MLP（保留参考）\n│   ├── rollout/\n│   │   ├── gpu_rollout_buffer.py   # CUDA-resident rollout buffer\n│   │   └── index_only_buffer.py    # 索引化观测（Phase 9+）\n│   ├── inference.py            # ONNX CPU 推理\n│   ├── onnx_export.py          # SB3 → ONNX 导出\n│   ├── price_limits.py         # 板别动态涨跌停\n│   ├── reward_functions.py     # Return / Sharpe / Sortino / Mean-Variance / MainWaveHold\n│   ├── main_wave_labels.py     # Phase 22 — MA5/MA10 死叉 + 5d cap 持仓回报标签\n│   ├── metrics.py              # 训练指标 JSONL 读写\n│   ├── wandb_integration.py    # 实验跟踪（默认离线）\n│   ├── sb3_callbacks.py        # SB3 callbacks (WandbMetricsCallback, GpuSamplerCallback, …)\n│   ├── gpu_monitor.py          # pynvml-based GPU 采样\n│   └── factors/                # polars-native 因子库\n│       ├── alpha101/            # WorldQuant Alpha101 (105 因子，10 模块)\n│       ├── gtja191/             # 国泰君安 Alpha191 (191 因子，10 batch 文件)\n│       ├── _ops.py              # 25+ 通用算子（ts_sum / ts_corr / cs_rank / decay_linear / regbeta / ...)\n│       ├── registry.py          # ALPHA101_REGISTRY + GTJA191_REGISTRY\n│       └── _docs.py             # markdown 文档生成器\n├── scripts/\n│   ├── train.py                # 训练入口 V1（CLI）\n│   ├── train_v2.py             # 训练入口 V2（Phase 21+ Dict obs，被 Phase 22 V1 main_wave 回退后保留）\n│   ├── infer.py                # 推理入口（CLI）\n│   ├── eval_backtest.py        # 测试集 IC / Sharpe / 等权净值曲线\n│   ├── _eval_main_wave_v1.py   # Phase 22 V1 main_wave 评估（含 hold_return / win_rate / drawdown）\n│   ├── compare_rewards.py      # 多 reward 类型对比训练\n│   ├── export_factor_panel.py  # PG → Parquet 数据抽取（含 SQL 模板）\n│   ├── generate_synthetic.py   # 合成 demo 数据生成\n│   ├── oss_download_resumable.py  # HEAD + Range-based resumable downloader\n│   └── reference_data/         # alpha101 / gtja191 reference parquet 重建脚本\n├── web/                        # Next.js 16 dashboard（runs/ 可视化）\n├── data/\n│   ├── README.md               # 数据格式 + 列名约定\n│   └── synthetic_demo.parquet  # 10 MB 开箱即跑\n├── docs/\n│   ├── ARCHITECTURE.md\n│   ├── FACTORS.md              # 因子前缀约定 + 列名规范\n│   ├── TRAINING_HISTORY.md     # 14 phase 完整训练栈演化（1350 行）\n│   ├── factor_library/         # 296 篇因子 markdown 文档\n│   ├── phase26/                # Phase 26A-G 实验报告\n│   ├── SCHEMA.md\n│   ├── TRAINING.md\n│   └── INFERENCE.md\n├── tests/                      # 1386+ 测试，含因子 parity 与 docs 验证\n├── handoffs/                   # 跨机器（4070 ↔ ECS）交接日志\n└── examples/\n    └── quickstart.py           # 端到端示例\n```\n\n**English.** Repository layout (see Chinese tree above for full structure). Three key code surfaces:\n\n- `src/aurumq_rl/env.py` and `gpu_env.py` — the Gymnasium `StockPickingEnv` and its later GPU-vectorized counterpart\n- `src/aurumq_rl/policies/per_stock_encoder.py` — the Deep Sets-style permutation-equivariant policy that became the architectural breakthrough in Phase 5\n- `src/aurumq_rl/main_wave_labels.py` — Phase 22's MA5/MA10 death-cross + 5d-cap hold-return reward that broke the 5.72 % random baseline for the first time\n\n### 2.5 硬件与训练资源约束 / Hardware \u0026 Training-Resource Constraints\n\n**中文.** 项目对硬件有两条**红线**：\n\n1. **本地 ECS（8C14G）严禁运行训练**。PyTorch 安装即占 ~3 GB RSS，训练时 OOM 必杀；7-worker `ProcessPoolExecutor` 曾把整台主机 OOM-killed + 重启。训练只能在 GPU 实例（推荐本地 RTX 4070+ 或云端 RTX 4090 / A10 / V100）。\n2. **`max_workers=3` 是硬上限**（对所有 `ProcessPoolExecutor` / `ThreadPoolExecutor`），PostgreSQL `shared_buffers=2GB`，内存余量 \u003c 4 GB 时 PG 会被 OOM。\n\n实证训练成本（i7-13700K + RTX 4070 12 GB + 64 GB DDR5）：\n\n| 配置 | 因子数 | 训练步数 | wall time | 备注 |\n|---|---:|---:|---|---|\n| smoke (Phase 0) | 16 | 1k | 90s | 合成数据，CPU 即可 |\n| Phase 1 | 16 | 100k | ~50 min | alpha101 short panel，n_envs=8, fps 333 |\n| Phase 7 | 64 | 50k | ~7 min | GPU framework smoke, fps 1490 |\n| Phase 10 | 64 | 1M | ~8h（隔夜）| LayerNorm + dual pooling + bf16, fps 326 |\n| Phase 14 | 64 | 1M | ~6h（隔夜）| TF32 + unique-date, fps 460 |\n| Phase 16a (prod) | 343 | 300k | ~5h | 6 seeds 并行外推 |\n| Phase 22 (main_wave) | 343 | 300k | ~8h | 3-run 隔夜对照 (A/B/C) |\n| Phase 26F-v3 (prod) | 361 | 300k | ~5h | 3 seeds × 1 config |\n\n**English.** Two hard rules:\n\n1. **The local 8-core 14 GB ECS is FORBIDDEN for training.** PyTorch installation alone occupies ~3 GB RSS; training will OOM-kill the host. A 7-worker `ProcessPoolExecutor` once OOM-killed and rebooted the box. Train only on a GPU instance (local RTX 4070+ or cloud RTX 4090 / A10 / V100).\n2. **`max_workers=3` is a hard ceiling** for all `ProcessPoolExecutor` / `ThreadPoolExecutor`. PostgreSQL `shared_buffers=2GB`; PG OOMs when host free RAM \u003c 4 GB.\n\nMeasured training cost on i7-13700K + RTX 4070 12 GB + 64 GB DDR5 (see Chinese table above for the 8-config breakdown spanning smoke runs to overnight 1M-step phases).\n\n---\n\n## §3 因子库 \u003ca id=\"3-因子库\"\u003e\u003c/a\u003e\n## §3 Factor Library \u003ca id=\"3-factor-library\"\u003e\u003c/a\u003e\n\n### 3.1 自带因子计算引擎 / The Built-in Factor Engine\n\n**中文.** `src/aurumq_rl/factors/` 是 polars-native 实现的 296 个量价因子（105 alpha101 + 191 gtja191）。每个因子一篇 markdown 文档在 `docs/factor_library/`，含原文公式 + Polars 实现说明 + 引用。\n\n| Family | 实现 | quality_flag=0 (clean) | =1 (best-effort) | =2 (stub) |\n|---|---|---|---|---|\n| alpha101 | 101/101 + 6 自定义 | 88 | 13 | 0 |\n| gtja191 | 191/191 | 177 | 12 | 2 |\n\n`quality_flag` 语义：\n- **0 (clean)**：完整 + 数值稳定 + 跨平台 parity（如 alpha_001、gtja_159）\n- **1 (best-effort)**：实现合理但存在已知边界情况（如 alpha_017 在窗口=2 时 std=0 触发 NaN，已用 `fill_null` 处理但未触发 inf）\n- **2 (stub)**：实现存在但等价于占位（如 gtja_115/189 没有可靠数据源对应的 sd_pe_ttm 字段）\n\n注册表用法：\n\n```python\nimport aurumq_rl.factors.alpha101  # registers 107\nimport aurumq_rl.factors.gtja191   # registers 191\nfrom aurumq_rl.factors.registry import ALPHA101_REGISTRY, GTJA191_REGISTRY\n\npanel = pl.read_parquet(\"ohlcv.parquet\")  # 需含 OHLCV + vwap + amount\ndf = panel.with_columns([fn(panel).alias(name) for name, fn in ALPHA101_REGISTRY.items()])\n```\n\n通用算子 `_ops.py` 提供 25+ 个跨家族复用的算子：`ts_sum / ts_corr / cs_rank / decay_linear / regbeta / ts_argmax / ts_argmin / ts_min / ts_max / ts_rank / ts_delta / ts_delay / ts_std / ts_skew / ts_kurt / ind_neutralize / scale / signed_power / sign`，所有都是 polars expr-aware，可在懒求值 graph 里组合。\n\n**English.** `src/aurumq_rl/factors/` ships 296 polars-native price-volume factors (105 alpha101 + 191 gtja191), one markdown doc per factor under `docs/factor_library/` with original formula + Polars implementation notes + references. `quality_flag ∈ {0 clean, 1 best-effort, 2 stub}` per the Chinese table above. The common-operator module `_ops.py` provides 25+ Polars-aware operators (`ts_sum`, `ts_corr`, `cs_rank`, `decay_linear`, `regbeta`, …) that compose lazily.\n\n### 3.2 A 股私有因子族 / Private A-share Factor Families\n\n**中文.** 这是项目 11 个 A 股私有因子族，必须从用户自己的数据仓库算好后写进 Parquet。它们对应的中国市场原始数据源在欧美市场没有等价物：\n\n#### 3.2.1 `mf_*` 主力资金流速 (Money Flow Velocity, 14 + 6 cols)\n\n由用户上游 `scripts/compute_mf_panel.py` 输出，14 个基础列 + 6 个 `_log` 变体。例：\n\n- `mf_net_{1d,3d,5d,10d,20d,60d}` — 主力净流入累计（元）\n- `mf_buy_share_main` — 主力买入占比（**SHAP rank 7**，Path 4 模型里第 7 重要的特征）\n- `mf_net_accel_5_20` — 5d / 20d 流入加速度\n- `mf_net_5d_amount_ratio` — 5d 净流入 / 5d 成交额\n- `mf_net_{1d,3d,5d,10d,20d,60d}_log` — sign-preserving log1p 变体（2026-05-08 加入），公式 `sign(x) · ln(1 + |x|/total_amount)`，把原始 std=1.5×10⁸ 的\"元\"量级压到 std=0.040 的\"无量纲\"量级\n\n**HUGE_TAIL 事故**：原始 `mf_net_*d` 标准差从 1e8 到 1e10，跨 ts_code 的尺度差异极大（一只大盘股一天净流入 10 亿元，一只小盘股 1 万元）。Phase 24 在 `data_loader._cross_section_zscore` 里 z-score 之后仍然有量级，因为 polars 默认 `ddof=1` 在 3000-stock 截面下分母被极端值拉爆。修复：上游加 `_log` 变体后训练直接吃压缩量级，跨年泛化恢复。\n\n#### 3.2.2 `mfp_*` 主力筹码持仓 (Main Force Position, 12 cols)\n\n由 `src/aurumq/factors/main_force.py` 输出，与 `mf_*` **互补但完全独立**：\n\n- `mfp_elg_buy_ratio_20d` — 超大单买入占比 20 日\n- `mfp_lg_buy_ratio_20d` — 大单买入占比 20 日\n- `mfp_main_net_cum_pct` — 主力净流入累计百分位\n- `mfp_main_net_volatility_20d` — 主力净流入波动 20 日\n\n**Phase 16 关键事故**：`mfp_` 前缀曾被静默从 `aurumq_rl.data_loader.FACTOR_COL_PREFIXES` 漏掉，12 列输入完全没进训练。修复后 16a 跑出 adj Sharpe +1.593（vs 之前 plateau +1.165）。教训：**前缀注册表是 single source of truth，prefix-glob 漏一个前缀等于静默丢一族特征**。\n\n#### 3.2.3 `cyq_*` 筹码分布 (Chip Distribution, 3 cols)\n\n源自 Tushare cyq_perf 表，3 列：\n- `cyq_winning_ratio` — 当前价位上的获利筹码占比\n- `cyq_concentration_70` — 70% 筹码集中在多少价位区间\n- `cyq_cost_distance` — 当前价 vs 平均成本距离\n\n**Phase 26 关键事故**：cyq 是 A 股独有的因子（券商内部模型 + Tushare 加工），但 Tushare 历史只能回填到 2025-10-20，更早数据是 cyq_perf v1.0 用合成方法补的。结果：训练集（含合成回填）`cyq_cost_distance` std = 0.197；OOS 集（全部实采）std = 0.066，**3× 压缩**。跨截面 z-score 不抹平这种 mid-stream regime shift —— 模型学到的是合成数据的尺度，到真实数据上全错。Phase 26C2 切换到 v1.2 修正 cyq（bulk API 重新回填）后，T-1 lift 从 1.47× 反弹到 **2.61×**（甚至超过原始 23A 的 2.38×，且收敛快 4 倍）。\n\n#### 3.2.4 `hm_*` 主流游资席位 (Hot Money Seats, 6 cols)\n\n源自 Tushare `top_list` + `top_inst`：\n- `hm_net_5d` / `hm_net_20d` / `hm_net_60d` — 游资席位累计净买入\n- `hm_recent_active` / `hm_seat_count_30d` / `hm_top3_concentration` — 活跃席位 / 30 日席位数 / 前 3 名集中度\n\n**结构性 hard wall**：Tushare 龙虎榜数据 ≥ 2023-08-16 才存在，2018-2023.8 的 `hm_*` 永远是 NULL。Phase 20 长 panel 训练时 LightGBM 的 `use_missing=True` 自动处理，但 RL 训练时观测向量必须填 0（不能 NaN）。修复：`data_loader._fill_missing_with_zero_track_mask` 同时填 0 + 写 mask，模型可选择性地学到\"这一列 mask=1 时无效\"。\n\n#### 3.2.5 `hk_*` 北向资金 (Northbound, 4 cols)\n\n`hk_hold_chg_60d` 等，**SHAP rank 16**。结构性 null：港股通名单外的 25% A 股永远是 NULL。同样 `_fill_missing_with_zero_track_mask` 处理。\n\n#### 3.2.6 `inst_*` 机构持仓 (Institutional, 3 cols)\n\n`inst_appear_count_60d` / `inst_net_30d` — 龙虎榜机构席位活跃度。\n\n#### 3.2.7 `mg_*` 融资融券 (Margin, 3 cols)\n\n`mg_short_chg_20d` / `mg_balance_pct` 等。78% A 股有融资融券覆盖。\n\n#### 3.2.8 `senti_*` 情绪 (Sentiment, 3 cols)\n\n涨停板池 + 同花顺热度榜派生。**已知问题**：`senti_ths_hot_pct` 99% null，因为同花顺热度榜只追踪 ~3000 只热门股；非热门股在 2024-08-29 之前完全没有数据。Phase 26 数据质量审计把 `senti_ths_hot_pct` 列入 `include_columns_v1_clean.txt` 的永久排除清单。\n\n#### 3.2.9 `sh_*` 股东结构 (Shareholder, 2 cols)\n\n`sh_holder_num_chg_30d` — 股东户数变化。86% null（季度披露，日频面板上稀疏）。\n\n#### 3.2.10 `fund_*` 基本面 (Fundamentals, 4 cols)\n\n`fund_pe_ttm` / `fund_pb` / `fund_roe_ttm` / `fund_revenue_growth`。**SHAP rank 11 (pe_ttm) / rank 25 (roe_ttm)**。\n\n事故：688*** 科创板的 fund_pe_ttm 在 2025-08 之前缺失约 600 只 × 每天的 hole，因为 Tushare daily_basic 接口对科创板支持不全。2026-05-08 批次用 bulk daily_basic 回填了所有 600+ × 历史日期。\n\n#### 3.2.11 `ind_*` 申万行业相对强度 (Industry, 2 cols)\n\n`ind_relative_strength_20d` / `ind_relative_strength_60d` — 个股 vs 申万一级行业指数收益差。49-57% null（`sw_index_member` 表只覆盖约 3000 只主板成分股）。\n\n#### 3.2.12 `mkt_*` 大盘 (Market, 2 cols)\n\n`mkt_index_pct_chg_5d` / `mkt_index_volatility_20d` — 上证指数派生。**Phase 16 关键发现**：drop `mkt_*` 组反而 +0.428 lift！原因：`mkt_*` 在主板宇宙下高度共线（所有股票同一个上证指数派生量），模型用它做\"今天大盘涨/跌\"的偷懒预测，反而损害了选股能力。**永久从主流配置移除**。\n\n#### 3.2.13 `tech_*` / `cmf_*` / `zt_*` 技术指标 (Tech panel, 30 + 2 + 6 cols)\n\nPhase 26 新增。`tech_*` 30 列 = MA5/10/20/60 比值 + KDJ 派生 + MACD 派生 + Bollinger 派生 + ATR 派生 + 振幅。`cmf_*` 2 列 = Chaikin Money Flow 60d/120d。`zt_*` 6 列 = 涨停板 30d/60d 频次 + 首板/连板/最长连板。\n\n**Phase 24/25 重大事故**：24A 把 36 个技术因子直接接在 RL 训练 panel-load 时算（而不是上游 parquet），结果 T-1 hit 从 2.11% 跌到 **0.40%**（lift 0.45× \u003c 随机 0.89×）。根因：KDJ/振幅近似自 close-only（panel 没有 OHLC），MA-cross / golden-cross 是二值事件标志，进 LayerNorm 后 z-score 把 binary 0/1 拉成极端 outlier，污染了梯度。Phase 26F 修复：把二值事件改为 **指数衰减 τ=10d 编码** `evt(t) = sum(1[event in last 10d] * exp(-(t-tau)/10))`，T-1 hit 从 1.13× 反弹到 **2.27×（best 2.41% hit at step 50k）**。教训见 §7。\n\n**English summary.** The 11 private A-share factor families are: `mf_*` (Money Flow Velocity, 14 base + 6 sign-preserving log variants — fixes the 1e8-yuan HUGE_TAIL scale issue); `mfp_*` (Main Force Position, 12 cols, independent of `mf_*` despite the similar prefix — Phase 16 found `mfp_` was silently missing from `FACTOR_COL_PREFIXES`); `cyq_*` (chip distribution, 3 cols — Phase 26C2 v1.2 fix recovered T-1 lift from 1.47× to 2.61× by replacing the synthetic-backfill v1.0 with a bulk-API-recomputed v1.2); `hm_*` (Dragon-Tiger hot-money seats, 6 cols, structural null pre-2023-08-16); `hk_*` (Northbound, 4 cols, structural null for 25 % non-HK-Stock-Connect stocks); `inst_*` (institutional, 3); `mg_*` (margin trading, 3); `senti_*` (sentiment, 3, 99 % null for non-hot stocks); `sh_*` (shareholder, 2, 86 % null due to quarterly disclosure); `fund_*` (PE/PB/ROE/revenue growth, 4 — SHAP rank 11/25); `ind_*` (SW industry relative strength, 2); `mkt_*` (market index, 2 — **dropped permanently in Phase 16 because removing it gave +0.428 adj Sharpe lift**). Phase 26 added `tech_*` (30), `cmf_*` (2), `zt_*` (6) — note that binary event flags must be exp-decay encoded (`τ=10d`) not raw 0/1 to avoid the −33 % regression seen in Phase 24A.\n\n### 3.3 因子前缀注册纪律 / Factor-Prefix Registration Discipline\n\n**中文.** `aurumq_rl.data_loader.FACTOR_COL_PREFIXES` 是 single source of truth。当前规范清单：\n\n```python\nFACTOR_COL_PREFIXES = (\n    \"alpha_\", \"mf_\", \"mfp_\", \"hm_\", \"hk_\", \"inst_\", \"mg_\", \"senti_\",\n    \"sh_\", \"fund_\", \"ind_\", \"cyq_\", \"gtja_\",\n    \"tech_\", \"cmf_\", \"zt_\",\n)\n```\n\n**漏一个前缀 = 静默丢一族特征 + Phase 16 复现**。所有 PR 修改这个 tuple 必须同步更新：\n\n- `tests/test_data_loader.py:test_factor_col_prefixes_lockdown` —— 字典对比 + 顺序对比\n- `scripts/export_factor_panel.py:FACTOR_PREFIXES` —— PG 抽取脚本的镜像列表\n- `docs/FACTORS.md` —— 表格 + 列名规范文档\n\n**English.** `aurumq_rl.data_loader.FACTOR_COL_PREFIXES` is the single source of truth (17 prefixes today). Missing one = silently lose a factor family = Phase 16 reproduction. Three sites must be kept in sync per PR: the tuple itself, the `tests/test_data_loader.py` lockdown test, and `scripts/export_factor_panel.py`.\n\n### 3.4 SHAP 剪枝实验：345 → 226 / SHAP-Based Pruning\n\n**中文.** 见 [`handoffs/2026-05-10-sl-extras/shap_audit/`](handoffs/2026-05-10-sl-extras/shap_audit/)（paris 侧执行）：\n\n- **方法**：`shap.TreeExplainer` 跑在 Path 4 最佳 LightGBM 单模（`nl31_lr050_mdl50_seed44`），10k 行 VAL_EFF 数据，按 `mean(|SHAP|)` 排名。\n- **Top 5 surprises**：`gtja_159` 一骑绝尘（mean|SHAP|=0.001270，gain 28.7%，162 个分裂点）、`gtja_158`、`gtja_065`、`gtja_140`、`gtja_181`。资金流第 7 名 `mf_buy_share_main`，基本面第 11 名 `fund_pe_ttm`，北向第 16 名 `hk_hold_chg_60d`。\n- **剪枝规则**：`mean(|SHAP|) \u003c 1e-6` 视为零贡献 → 119 个候选 → 保存到 `drop_candidates.json`。被剪掉的例子：`alpha_098`、`gtja_054`、`gtja_101`、`gtja_190`、`alpha_002`、`gtja_001`、`mf_net_accel_5_20`、`mf_net_60d`、`gtja_114`、`inst_net_30d`。\n- **Path 6 验证**（Bayesian opt 50 trials）：226 列训出来 H1 校准 mean_y = +0.028265 vs Path 4 全 345 列的 +0.028483，**Δ = −0.0002**（与噪声不可分辨）。bundle 从 40 MB 缩到 32 MB，训练时间 −15%。结论：**超参数搜索已 saturate，剪枝是免费午餐**。\n\n**English.** SHAP-based feature pruning ran on the best Path 4 LightGBM single model (`nl31_lr050_mdl50_seed44`) over 10k VAL_EFF rows. `mean(|SHAP|) \u003c 1e-6` ⇒ drop candidate ⇒ 119 columns saved to `drop_candidates.json`. Validating on Path 6 (226 cols + 50 Bayesian-opt trials): H1 calibrated mean_y +0.028265 vs full-345 Path 4 +0.028483 = −0.0002 (indistinguishable from noise). Bundle 40 MB → 32 MB, training −15 %. Lesson: **hyperparameter search is saturated; SHAP pruning is a free lunch.**\n\n### 3.5 存储路径与流式处理 / Storage Layout \u0026 Streaming\n\n**中文.** 因子 parquet 按家族 × 年份分片：\n\n| 路径 | 内容 |\n|---|---|\n| `data/duckdb/factor_eval/alpha_panel_year=YYYY.parquet` | alpha101 全族（109 cols） |\n| `data/duckdb/factor_eval/gtja_panel_year=YYYY.parquet` | gtja191 全族（193 cols，单年最大 1.37 GB） |\n| `data/duckdb/factor_eval/mf_panel_year=YYYY.parquet` | mf_ 22 cols (14 + 6 log + 2 helper) |\n| `data/duckdb/factor_eval/cyq_panel/year=YYYY.parquet` | canonical cyq 3 cols（v1.2 修正版） |\n| `data/duckdb/factor_eval/tech_panel/year=YYYY.parquet` | tech_ 30 cols |\n| `data/duckdb/factor_eval/tech_event_panel/year=YYYY.parquet` | tech_evt_ 8 cols（含 exp-decay 编码） |\n| `data/duckdb/quotes_enriched/year=YYYY.parquet` | 11 个内部 enriched 家族 `mfp_/hm_/hk_/inst_/mg_/senti_/sh_/fund_/ind_/mkt_/cyq_` legacy |\n\n**流式 concat 红线**：14 GB ECS 上禁止跑 `pl.concat(diagonal_relaxed) + sink_parquet` 拼 10 年面板，会 OOM-killed 主机。正确做法是先 shard，然后 `pl.scan_parquet([shards], missing_columns=\"insert\")` 流式扫，见 `scripts/build_combined_panel_safe.py`。\n\n**English.** Factor parquets are sharded by family × year (see Chinese table). **Streaming red line**: on the 14 GB ECS, `pl.concat(diagonal_relaxed) + sink_parquet` over 10 years of panel data will OOM-kill the host. Correct pattern: shard first, then `pl.scan_parquet([shards], missing_columns=\"insert\")` streaming scan. See `scripts/build_combined_panel_safe.py`.\n\n---\n\n## §4 训练栈演进史 — Phase 0 → 14 \u003ca id=\"4-训练栈演进史--phase-0--14\"\u003e\u003c/a\u003e\n## §4 Training-Stack Evolution — Phase 0 to 14 \u003ca id=\"4-training-stack-evolution--phase-0-to-14\"\u003e\u003c/a\u003e\n\n**中文.** 本节是 GPU 训练栈本身的工程史 —— 从最初 11% GPU 利用率到 1M-step 隔夜训练的所有 stack diff、bug、消融。**模型/数据/奖励的实验史在 §5（Phase 15-26）**。两个 phase 编号体系独立：本节 Phase 0-14 是「框架建设」，§5 Phase 15-26 是「在已建好的框架上跑模型实验」。\n\n完整的逐 phase 记录在 [`docs/TRAINING_HISTORY.md`](docs/TRAINING_HISTORY.md)（1350 行），本节是其压缩版。\n\n**English.** This section is the engineering history of the training stack itself — every stack diff, bug, and ablation from the initial 11 % GPU utilization to the 1M-step overnight training. **The modeling / data / reward experiments live in §5 (Phases 15–26).** The two numbering systems are independent: Phases 0–14 here are \"framework construction\"; §5 Phases 15–26 are \"model experiments on top of the built framework\".\n\nThe full per-phase record is in [`docs/TRAINING_HISTORY.md`](docs/TRAINING_HISTORY.md) (1350 lines). This section is its compressed version.\n\n### Phase 0 — 合成数据流水线打通 / Synthetic Pipeline Bring-up\n*(~pre-2026-04-29)*\n\n**Goal.** Prove the parquet → env → SB3 PPO → ONNX → backtest → JSON link before touching real data.\n\n**Stack.** `StockPickingEnv` (numpy panel, single-process), SB3 default `MlpPolicy net_arch=[64,64]`, PPO `n_envs=1 batch=64 n_steps=2048 epochs=10`, `synthetic_demo.parquet` (~200 SYN-coded fake stocks).\n\n**Bugs surfaced** (4 in PR #1 / #2):\n\n| # | Bug | Fix |\n|---|---|---|\n| 1 | `gymnasium` not always installed | lazy import + placeholder raises ImportError |\n| 2 | ONNX export device mismatch (CUDA policy + CPU dummy_obs) | move policy to CPU before export |\n| 3 | `torch.onnx dynamo=True` breaks SB3 `Normal` distribution | pass `dynamo=False` |\n| 4 | JSON serializer can't handle `numpy.float32` | `WandbMetricsCallback._append_jsonl` got `default=_json_default` |\n\n**Outcome.** Pipeline end-to-end. ONNX exported. Backtest IC ≈ 0 (synthetic noise, expected).\n\n### Phase 1 — 第一次真实数据训练 / First Real-Data Run *(2026-04-29 ~ 30)*\n\n**Goal.** Scale to a real factor panel and a real GPU. Run a 100k-step PPO on alpha101 to see how far naive setup goes.\n\n**Data.** `factor_panel_alpha101_short_2023_2026.parquet` — 105 alpha cols, 5743 stocks, 2023-01..2026-04. After `main_board_non_st`: 3043 × 800 × 105.\n\n**Config.** PPO `--total-timesteps 100000 --n-envs 8 --vec-normalize --learning-rate 3e-4 --target-kl 0.05 --max-grad-norm 0.3 net_arch=[64,64] n_factors=16`.\n\n**Bugs (8 new):** NaN propagating through cross-section z-score (real PG data has NaN cells for suspended / pre-IPO; synthetic didn't); OOS obs_dim mismatch (training universe = 3043, OOS = 3052 because some IPOs landed — env's `observation_space` is fixed at training time; fix = `align_panel_to_stock_list` persisted in `metadata.json`); PPO `approx_kl=41,820` on first update (12.5M-param first layer + 48,688-dim obs; fix = `target_kl=0.05 + max_grad_norm=0.3` → `approx_kl=0.028`); `mean_fps=0` in summary (SB3 only emits `time/fps` on rollout-summary frames; fix = callback computes wall-time fps); `metrics_summary` all-null (callback wrote raw SB3 keys; `summarize_metrics()` expects canonical schema; fix = raw→canonical mapping at write time); `runs/` gitignore unanchored (`web/app/runs/` silently dropped; fix = `/runs/` anchored at root); alpha045 STHSF parity 44 % mismatch on Windows only (scipy rank-tie-break unstable across versions on 10-stock synthetic; fix = `@pytest.mark.xfail(strict=False)`); OSS admin AK disabled mid-flight (switch to wepa AK).\n\n**Outcome.** 100k-step PPO ran clean. **fps ~ 333. GPU util ~ 11 %.** The 4070 was massively underutilized — wide first-layer of `[64,64]` was only 3 M params; GPU spent most of its time waiting on CPU rollouts.\n\n### Phase 2 — 联合 panel + 网络加宽 + feature_group_weights / Combined Panel + Network Widening *(2026-04-30)*\n\n**Data added.** `factor_panel_combined_short_2023_2026.parquet` — **355 factor cols** (105 alpha + 191 gtja + 14 mf + 12 mfp + 5 hk + 4 fund + 3 inst + 3 mg + 3 senti + 2 sh + 2 ind + 2 mkt + 6 hm + 3 cyq), 5643 stocks × 800 dates, 7.7 GB zstd. After main-board filter: 3014 × 600.\n\n**Code added.** `--policy-kwargs-json` CLI accepts `{\"net_arch\":[2048,1024,512], \"activation_fn\":\"relu\"}`. `--feature-group-weights-json` accepts e.g. `{\"alpha_*\":2.0, \"mf_*\":0.5}`, applied **after** z-score in `_apply_feature_group_weights` so `VecNormalize` doesn't neutralize the weights.\n\n**Network widening.** `net_arch=[2048,1024,512]`. First-layer params for `n_factors=64`: `3014 × 64 × 2048 ≈ 395 M`. GPU memory 3 GB → 12 GB peak; util peak 11 % → 57 %.\n\n**3-way alpha-prefix ablation:**\n\n| Run | `--feature-group-weights-json` | OOS IC | OOS top30 Sharpe |\n|---|---|---|---|\n| `ablation_alpha_w0_5` | `{\"alpha_*\":0.5}` | (~0) | (~random) |\n| `ablation_alpha_w1_0` | `{\"alpha_*\":1.0}` (no-op baseline) | (~0) | (~random) |\n| `ablation_alpha_w2_0` | `{\"alpha_*\":2.0}` | +0.0006 | −0.807 (random p50 −0.482) |\n\n**Decision.** Framework works end-to-end. Numbers are noise at 15k steps. Validation passed; promote `feature_group_weights` as load-bearing CLI feature.\n\n### Phase 3 — 三轮 smoke R1/R2/R3 / Three Smoke Rounds *(2026-05-01 morning)*\n\n**Round 1 (R1).** `--policy-kwargs-json '{\"net_arch\":[1024,512,256]}'` + `--feature-group-weights '{\"alpha_*\":2.0, \"mf_*\":1.5, \"gtja_*\":1.0}'` at 50k steps, `n_envs=12, n_steps=2048, target_kl=0.05`. **First model with explained_variance climbing.** `value_loss` from 1.5e-2 → 4.3e-3 over 22 rollouts. OOS IC = +0.011, top30 Sharpe = +1.42 (random p50 −0.48, vs-p50 +1.90). **GPU util 35 %, fps 312.**\n\n**Round 2 (R2).** Three changes at once: `target_kl 0.05 → 0.10`, `n_envs 12 → 14`, `n_steps 2048 → 4096`. First attempt OOMed (`MemoryError` allocating 8.83 GiB rollout buffer); reduced `n_steps 4096 → 1024`; ran again. **OOS top30 Sharpe = +2.16**, vs-p50 = +0.74 above R1. Convergence accelerated (explained_var 0.93 at 30k vs R1's 0.81). **But: three changes at once = uninterpretable**. Could be KL relaxation, env parallelism, or buffer length. Lesson recorded — Phase 3's central rule: **one change per round**.\n\n**Round 3 (R3).** Just `target_kl 0.10` + `learning_rate 3e-4 → 1e-4 anneal`. OOS top30 Sharpe = +1.89 (R2 = +2.16, R1 = +1.42). Within seed variance of R2; suggests `target_kl` accounts for most of R2's lift, but `n_envs / n_steps` cannot be cleanly attributed.\n\n**Lesson.** **OOS Sharpe at 50k steps is noise.** Don't pick winners from smokes; pick them from convergence-scale runs (≥ 1M ideally 5M). Burned three rounds arguing about R1/R2/R3 ranking before admitting differences were within seed variance.\n\n### Phase 4 — fps 扩展实验 + IPC 天花板 / fps Scaling and IPC Ceiling *(2026-05-01 noon)*\n\n**Goal.** Find the n_envs ceiling.\n\n**Method.** Smoke-grid `n_envs ∈ {12, 14, 16, 18, 20}` at fixed `n_steps=1024`.\n\n| `n_envs` | fps | GPU util | Outcome |\n|---:|---:|---:|---|\n| 12 | 314 | 35 % | R1 baseline |\n| 14 | 366 | 41 % | linear scale |\n| 16 | 412 | 47 % | linear scale |\n| 18 | 455 | 53 % | linear scale starting to bend |\n| 20 | 458 | 56 % | **bent — IPC bottleneck** |\n\n**Realization.** Above `n_envs=18`, adding env doesn't proportionally raise fps because the bottleneck is **Python IPC** between worker subprocs and the central learner, not GPU compute. **And: n_envs=20 OOMed** on rollout buffer alloc (14.7 GiB). Back to n_envs=12 for safety.\n\n**The IPC ceiling discovery** changed the project direction. The classic SB3 setup (numpy panel + subproc envs + CPU rollouts → GPU train) is **fundamentally CPU-rollout-bound**. The 4070 was sitting at 56 % util at the bottleneck. To break through, we'd need to move rollouts onto the GPU itself. → Phase 5 / 6.\n\n### Phase 5 — 四个 realization 推出 GPU-框架重构 / Four Realizations Driving GPU-Framework Redesign *(2026-05-01 afternoon)*\n\n**Goal.** Sit down, look at the data, decide whether to keep tuning or fundamentally redesign.\n\n**Four realizations:**\n\n1. **Brute-force capacity is a trap.** 395 M-param flat MLP needed 12 GB VRAM, fps capped at 458, and didn't beat the per-stock symmetry prior. A symmetry-correct architecture (per-stock encoder, ~50 k params shared across all stocks) has dramatically more inductive bias for stock-picking AND is faster.\n2. **The numpy panel is the wrong abstraction.** Re-uploading the same panel to GPU memory every env reset is wasteful. The panel should live in GPU memory throughout training, indexed by env step.\n3. **Observations should be indices, not values.** Stock factor vectors don't change across env steps — only which date the env is at changes. Send (date_idx, stock_codes_idx) over the IPC boundary, do the GPU-side gather afterwards.\n4. **VRAM and RAM are different.** Confused them once: GPU showed 12 GB used, my Python proc was only 3.8 GB RSS. Always check `nvidia-smi --query-gpu=memory.used` per-process AND host `Get-Process -RSS`.\n\n**Decision.** Design a GPU-vectorized framework: per-stock encoder + CUDA-resident panel + index-only observations + dual pooling head. Implementation in Phase 6/7.\n\n### Phase 6 / 7 — GPU-vectorized 框架 + 50k smoke / GPU Framework + 50k Smoke *(2026-05-01 evening)*\n\n**Design (Phase 6, designed afternoon; built in Phase 7).**\n\n- **`gpu_env.py`**: Gymnasium-compatible env that holds the entire panel as a CUDA tensor `panel[n_dates, n_stocks, n_factors]`. Step = advance one date; obs = `panel[date_idx]` slice + holdings mask. **n_envs=16 in a single proc** (vectorized across env axis on GPU, no IPC).\n- **`PerStockEncoderPolicy`** (Deep Sets): apply the SAME MLP to each stock's factor row, then aggregate with mean+max dual pooling. Permutation-equivariant. Net_arch=[64, 32] per-stock, then a [64, 32] head — only ~50 k params total, shared across 3014 stocks.\n- **Action**: scaled tanh on per-stock logits, top-K selection.\n\n**Smoke run (Phase 7).** 50k steps, n_envs=16, `n_steps=1024`. **fps 1490** (vs Phase 4 ceiling 458 — 3.25× lift). GPU util peak 78 %, mean 62 %. OOS IC +0.014, top30 Sharpe +1.78 — better than Phase 3 R3 +1.89 was within noise of, and **convergence reached at 30k vs R3's 50k**.\n\n**Result.** The GPU framework eclipses the previous best at one-third the wall time. The redesign was worth it.\n\n### Phase 8 — GPURolloutBuffer (CUDA-resident) *(2026-05-01 evening)*\n\n**Bottleneck.** Even with GPU env, rollout buffer was numpy on host. Every PPO update did host→device copies for each minibatch.\n\n**Fix.** Wrote `gpu_rollout_buffer.py`: holds `obs / actions / values / log_probs / rewards / advantages / returns` as CUDA tensors. PPO update reads directly from device memory — zero copies.\n\n**Outcome.** fps 1490 → 1820. GPU util mean 62 % → 71 %. VRAM +1.2 GB (acceptable).\n\n### Phase 9 — IndexOnlyRolloutBuffer + n_steps=1024 / Index-Only Observations *(2026-05-01 late evening)*\n\n**Bottleneck identified.** Even with cuda-resident buffer, each entry stored `obs = (n_stocks × n_factors)` float32 = 3014 × 64 × 4 = 770 KB per step. At `n_envs=16, n_steps=1024`: 16 × 1024 × 770 KB = **12.6 GiB** rollout buffer. We were paying for storing the entire factor cross-section in memory N_envs × N_steps times.\n\n**Insight.** All observations are slices of the same panel. Store **only the date index** (4 bytes) and gather on-the-fly during minibatch.\n\n**Fix.** `IndexOnlyRolloutBuffer`: stores `date_idx[n_envs, n_steps] int32` (~64 KB) + `holdings_mask[...] bool` (~250 KB). Gather `panel[date_idx[batch]]` at minibatch read time. Effective batch size keeps the same numerical behavior; memory is **200× smaller**.\n\n**Outcome.** Rollout buffer 12.6 GiB → 0.06 GiB. **Freed VRAM lets us raise `n_envs=16 → 32` and `n_steps=1024 → 2048`** within the same 12 GB budget. fps 1820 → 2050. GPU util peak 88 %.\n\n### Phase 10 — Optimizer-Orphan Bug + LayerNorm + Dual Pooling *(2026-05-01 night)*\n\n**Bug.** `value_loss` plateaued around 4e-3, never broke below. Suspected vanishing gradients in the value head. Inspected: **value head's parameters were not in the optimizer**. SB3 default uses one optimizer for both policy and value when they share an extractor; my custom policy split them and only registered policy params.\n\n**Fix.** Explicit `optim.AdamW([{params: policy.parameters()}, {params: value_net.parameters(), lr: 1e-3}], lr=3e-4)`. Value head learning unlocked. `explained_variance` climbed 0.78 → **0.99** within 50k steps.\n\n**Two more additions in this phase:**\n\n- **LayerNorm after each per-stock MLP layer.** Real-data cross-section z-scores still have outliers (after `nan_to_num`, a single inf cell at the cell level can still pull mean). LayerNorm gives stable gradients. ~Verified ablation: removing LayerNorm gave value_loss instability across seeds.\n- **Dual pooling head**: aggregate per-stock representations by `concat(mean, max)` instead of pure mean. Mean captures average market state; max captures the most-extreme stock signal. Worth +0.06 explained_var on the 50k smoke.\n\n**Outcome.** fps 2050 → 1980 (the LayerNorm + dual pool cost ~3 %, but `explained_var=0.99` was worth it). bf16 was attempted but eliminated in Phase 11.\n\n### Phase 11 / 12 — bf16 / adaptive target_kl (eliminated) *(2026-05-01 night)*\n\n**Phase 11.** Tried `torch.autocast(dtype=bfloat16)` for matmul+linear. Memory −20 %, fps +12 %. **But**: `approx_kl` became unstable, occasionally spiking to 0.3 (vs nominal 0.02). Inspected: tail of policy logits at bf16 dynamic range had quantization noise that compounded in KL computation. **Rolled back to fp32**.\n\n**Phase 12.** Tried `target_kl=0.10` adaptive (raise to 0.15 if violated 3x in a row). PPO update frequency dropped from every rollout to ~70 %. Total update count similar, value loss similar, IC similar — **no measurable benefit**. **Removed for code simplicity.**\n\n### Phase 13 — PPO SGD perf-probe / Profiling *(2026-05-01 late night)*\n\n**Goal.** Profile a single PPO update with torch.profiler to find any remaining low-hanging fruit.\n\n**Findings.**\n- 53 % of SGD time was in advantage computation (`compute_returns_and_advantage`).\n- 22 % in policy log-prob recomputation.\n- 11 % in value head forward.\n- 14 % in `optimizer.step()`.\n\n**Fix.** Wrote `compute_returns_and_advantage_vectorized()` that uses prefix-sum on CUDA tensors instead of Python for-loop over time steps. **53 % → 7 %**. SGD wall time per update −40 %.\n\n**fps 1980 → 2580.** GPU util mean 78 %.\n\n### Phase 14 — TF32 + unique-date + 1M overnight / TF32 + Unique-Date + 1M Overnight *(2026-05-02 → 2026-05-03 early)*\n\n**Goal.** Capacity build. Run 1M steps overnight on the post-Phase 13 stack to test stability and final convergence.\n\n**Two micro-optimizations** going in:\n\n- **TF32**: `torch.backends.cuda.matmul.allow_tf32 = True; cudnn.allow_tf32 = True`. Free ~12 % matmul speedup on RTX 4070 (Ampere) with no measurable accuracy loss.\n- **Unique-date dedup**: same date often appears multiple times in a 16-env × 2048-step rollout because envs reset asynchronously. Detected and gathered only unique `date_idx` once per minibatch, then duplicated rows. Saves ~30 % gather time.\n\n**1M-step overnight run** (alpha101+gtja191 296-col on short panel; n_envs=16, n_steps=2048, target_kl=0.05 fixed, lr=3e-4 → 1e-5 linear, ent_coef 0.01 → 0):\n\n- Wall time: **5h 47min**\n- fps mean: **460** (started 326 from cold start, reached 540 at ~200k)\n- GPU util mean: 73 %\n- Peak VRAM: 9.8 GiB / 12 GiB\n- approx_kl trajectory: stable 0.025-0.045, no spikes\n- explained_variance: 0.99 from ~150k onwards\n- Final OOS top30 Sharpe: **+5.83** (vs random p50 +0.62, vs-p50 +5.21)\n\n**Outcome.** Stack hardened. Phase 14 is the \"ready for real experiments\" milestone. All Phase 15+ modeling experiments run on this exact stack.\n\n**Section recap.** From Phase 0's 1k-step smoke (fps 50ish) to Phase 14's 1M-step overnight (fps 460), the framework went through **15 stack diffs, 24 documented bugs, and 5 mid-flight algorithmic redesigns**. The biggest single jump was Phase 5/6/7 — moving rollouts onto the GPU **3.25×'d fps**, and the symmetry-correct per-stock encoder simultaneously **reduced parameter count by 8000×** (395 M flat MLP → 50 k Deep Sets) while improving OOS Sharpe.\n\n---\n\n## §5 模型实验史 — Phase 15 → 26 \u003ca id=\"5-模型实验史--phase-15--26\"\u003e\u003c/a\u003e\n## §5 Modeling Experiment History — Phase 15 to 26 \u003ca id=\"5-modeling-experiment-history--phase-15-to-26\"\u003e\u003c/a\u003e\n\n**中文.** Phase 14 之后框架不再变了。Phase 15-26 是**在固定 stack 上跑模型/数据/奖励实验**。每个 phase 都有明确假设、单变量改动、量化决策。\n\n**English.** After Phase 14 the framework stopped changing. Phases 15–26 are **model / data / reward experiments on a frozen stack**, each with a clear hypothesis, single-variable change, and quantitative decision.\n\n### Phase 15 — RL serving 集成 / RL Serving Integration *(2026-05-02 ~ 03)*\n\n**Goal.** Take the best 1M-step model from Phase 14 and integrate it into the AurumQ platform for live serving.\n\n**Three SB3 PPO models registered:**\n\n| Agent ID | OOS Sharpe | Note |\n|---|---:|---|\n| `phase15e_150k_grand_champion` | **+6.277** | active production model |\n| `phase15e_100k_alt_peak` | +5.94 | alternative early-peak ckpt |\n| `phase15e2_225k_continuation_peak` | +5.83 | continuation of Phase 15e but with extended steps |\n\n**Bundle layout** `models/rl/\u003cid\u003e/`:\n- `policy.zip` — SB3 native (kept for `PPO.load` path)\n- `factor_schema.json` — factor name list, ordering, hash\n- `metadata.json` — train_start_date, train_end_date, stock_codes, feature_group_weights, factor_count, policy_class\n- `golden_inference.json` + `golden_obs.npy` + `golden_scores.npy` — sanity-check pair persisted with the bundle\n- `checksums.json` — sha256 of all artifacts\n\n**5 admin-only endpoints** under `/api/v1/rl/agents/`:\n- `POST /import-bundle` — upload, validate schema, persist\n- `POST /{id}/validate` — replay `golden_obs` through policy, compare to `golden_scores`\n- `POST /{id}/archive` — soft delete\n- `POST /{id}/inference` — async job (returns job_id)\n- `GET /api/v1/rl/inference-jobs/{job_id}` — poll\n\n**Tech-debt addressed.** SB3 `PPO.load(device='cpu')` + LRU(3) policy cache + single-flight lock to prevent concurrent reloads under high traffic.\n\n**Frontend.** `ModelHashBadge`, `MissingDataAlert`, `RlAgentsView`, `RlStockPicksView`, `useInferenceJob` composable.\n\n### Phase 16 — 修复 eval bug 重新基线 / Eval Bug Fixes Force New Baseline *(2026-05-03, 4h)*\n\n**Goal.** Re-validate Phase 15's \"drop `mkt_*` group helps\" finding under three independent bug fixes.\n\n**Three bugs:**\n\n1. **Reward double-shift fix.** `FactorPanelLoader` was already encoding the fp-day forward return, but the env AND the importance-permutation pass were re-indexing `t+fp` ⇒ rewards were `fp` days too late.\n2. **Sharpe over-annualization.** Overlapping fp-day forward returns must be annualized by `√(252/fp)`, not `√252`. For fp=10, that's an inflation factor of ~3.16×. Phase 15's legacy Sharpe +6.277 includes this inflation; the **bug-corrected** Sharpe is ~+1.99 on the same data.\n3. **`mfp_` prefix silently missing.** 12 mfp cols had been dropped from `FACTOR_COL_PREFIXES`. Training input was 343 cols, not 355.\n\n**Models retrained.** 16a (drop `mkt_` only), 16b (drop `mkt_+gtja_`), 16c (extend 16b to 450k).\n\n**Key findings (under bug-fixed eval):**\n\n| Run | adj Sharpe | vs random p50 | IC | Note |\n|---|---:|---:|---:|---|\n| 15e legacy (uncorrected) | +6.27 | — | — | annualization artifact |\n| **16a (drop mkt_)** | **+1.593** | **+0.428** | +0.0143 | new prod candidate |\n| 16b (drop mkt_+gtja_) | +1.32 | −0.27 | +0.0109 | gtja_ is load-bearing, contrary to Phase 15 belief |\n| 16c (16b @ 450k) | +1.36 | −0.23 | +0.0112 | extension didn't help |\n\n**Two robustly anti-helpful groups emerged** in permutation importance: cyq (−0.142 ± 0.044) and inst (−0.115 ± 0.030). mfp turned out weakly positive (+0.047 ± 0.067), gtja_ load-bearing (+0.160 ± 0.126).\n\n**Decision.** 16a → production. Phase 15 legacy peaks retired as annualization artifacts. Phase 17 scoped to (a) seed-sweep 16a (b) test \"drop cyq+inst\" hypothesis.\n\n### Phase 17 — 种子鲁棒性 + 条件重要性陷阱 / Seed Robustness + The Conditional-Importance Trap *(2026-05-03, 7h)*\n\n**Goal.** Test whether Phase 16's \"robust anti-helpful cyq+inst\" signal transfers under retrain; measure seed dispersion.\n\n**Method.** 17A: train drop_mkt+cyq+inst at seed=42, 300k. 17B/C/D: re-run 16a at seeds 1/2/3. 17E: extend 16a (seed=42) to 450k.\n\n**Key findings.**\n\n- **17A failed catastrophically.** adj Sharpe = +0.861, vs-p50 = **−0.304** (16a was +0.428). The cyq+inst drop hypothesis is **FALSE**. The \"robust negative permutation importance\" signal turned out to be **conditional on the trained policy, not causal**.\n- **Seed sweep.** 3/4 seeds beat random; mean lift +0.249; range [−0.060, +0.428]. seed=42 sits at the upper edge of the noise band; seed=2 is a lone failure.\n- **17E (450k) produced no new peak.** Phase 16a's +1.593 at step 224k is confirmed as the seed=42 global maximum.\n\n**Critical lesson.** **Stop chasing factor drops based on permutation importance alone.** Permutation importance reflects **what the trained policy uses**, not **what is causally helpful for prediction**. To test causality you must retrain after the drop.\n\n### Phase 18 — 多种子集成 / Multi-Seed Ensemble *(2026-05-03, 6h)*\n\n**Goal.** Convert the seed-sensitivity finding into a deployable mitigation.\n\n**Method.** Add seeds 4-7 (18A-D) at the unchanged 16a config. Build rank-mean / z-mean / z-median ensembles.\n\n**Key findings.**\n\n| Run | adj Sharpe | vs random p50 |\n|---|---:|---:|\n| seed=42 (16a) | +1.593 | +0.428 |\n| seed=1 (17B) | +0.97 | +0.115 |\n| seed=2 (17C) | +0.40 | −0.060 — failure |\n| seed=3 (17D) | +1.42 | +0.408 |\n| **seed=4 (18A)** | **+1.917** | **+0.752** — single-seed big win |\n| seed=5 (18B) | +1.84 | +0.596 |\n| seed=6 (18C) | +0.92 | +0.080 |\n| seed=7 (18D) | +0.27 | −0.140 — failure |\n\nAcross 8 seeds: mean vs-p50 +0.352, median +0.388, **win rate 6/8**. **Top-K Jaccard between seeds was 0.003–0.010** — seeds chose almost completely disjoint baskets. This is exactly why ensembling lifts.\n\n**The 6-member rank-mean ensemble** (excluding seeds 2 \u0026 7): vs-p50 **+0.711** (Δ vs seed=42 alone = +0.283), IC = **+0.0278** (1.94× 16a), non-overlap Sharpe +1.938.\n\n**Decision.** Ensemble passed strong-candidate gate. But: single-OOS-window optimum ≠ production Sharpe. Keep 16a live; advance ensemble as candidate pending fresh post-2026-04 holdout.\n\n### Phase 19 — 执行约束 + 多窗口验证 / Execution Constraints + Multi-Window Validation *(2026-05-03, 6h)*\n\n**Goal.** Stress-test `ens_rankmean6` against realistic costs, T+1 / limit-down filters, multi-window stability, and seed=4's contribution distribution.\n\n**Method.** Quarter blocks, rolling 60-day windows (step 20), execution simulation at 30/60/100 bps round-trip with limit-down deferral.\n\n**Key findings.**\n\n- **3/3 quarters won. 100 % rolling-60d win rate. 7/7 windows IC-positive.**\n- **Post-cost at 60 bps**: ensemble adj_S = **+0.971** vs 16a **+0.579** (Δ +0.392).\n- **At 100 bps**: gap **widened** to +0.272 vs −0.233 — **ensemble stayed positive when 16a flipped negative**.\n- **Seed=4 forensics warning**: **100.6 % of seed=4's marginal lift came from a single month (2026-01)**. Removing 16a from the ensemble actually improved score slightly (+0.045) — 16a was the weakest of the six.\n- **Fresh holdout check**: 0 dates past 2026-04-24 with fp=10, threshold ≥ 40 — **INSUFFICIENT**.\n\n**Decision.** Conclusion locked by data freshness. Phase 16a stays as live production; ensemble remains release-candidate. **No factor drops based on importance alone.** Phase 20 priority: collect fresh holdout.\n\n### Phase 20 — 长 panel 训练 / Long-Panel PPO *(2026-05-05)*\n\n**Goal.** Retrain 16a config on the long panel (2018-01 → 2025-06 train, 2025-07 → 2026-04 OOS) and check whether more history improves.\n\n**Two seeds:** 20A seed=42, 20B seed=4. Each 300k steps on the new 7y panel.\n\n**Key findings.**\n\n| Run | adj Sharpe | OOS hit_rate@5 | OOS win_rate |\n|---|---:|---:|---:|\n| Phase 16a (2y train) | +1.593 | 4.88 % | 36.9 % |\n| **Phase 20A (7y, seed=42)** | **+1.78** | 4.94 % | 37.4 % |\n| Phase 20B (7y, seed=4) | +1.42 | 4.85 % | 37.1 % |\n\n**Combined-evidence.** 2-seed average vs 16a: +0.012 adj Sharpe (within noise). The long panel is **at parity** with the short panel for the RL track on this metric.\n\n**20C cross-data ensemble: BLOCKED.** Ensembling the 7y-trained policy with the 2y-trained policy required obs_dim alignment, but the 7y policy was trained with 2018-2019 data that includes ~600 stocks that delisted before 2025; padding caused obs_dim mismatch. Decision: defer to Phase 21.\n\n**Decision.** Long panel doesn't help RL track (will revisit in Phase 22 after reward redesign). Phase 19's 16a stays live. The long-panel data **does** later become foundational for the SL track (see §6: P0 / Path 1 long / path5_long).\n\n### Phase 21 — V2 forward_10d 拒绝 / V2 forward_10d REJECTED *(2026-05-05 ~ 06)*\n\n**Goal.** Try a brand-new V2 architecture (Dict observation space + larger transformer-style head) on the existing forward_10d reward.\n\n**Result.**\n\n| Run | adj Sharpe | hit_rate@5 | win_rate | avg_hold |\n|---|---:|---:|---:|---:|\n| Phase 16a (V1) @ top_k=3 | +1.59 | 4.88 % | 36.9 % | +0.20 % |\n| **Phase 21A (V2 forward_10d)** | **−0.60** | **3.70 %** | **34.5 %** | **−0.16 %** |\n\n**Catastrophic regression.** V2 Dict obs scheme broke something about the policy's gradient flow; or maybe the Transformer attention head on 3014 stocks is too parameter-heavy for the 300k-step budget. Three retries with seed sweep — same result.\n\n**Decision.** **V2 rejected.** V1 `PerStockEncoderPolicy` stays canonical. Phase 22 reverts to V1 architecture but changes the reward function (see below).\n\n### Phase 22 — 主升浪奖励重设计 / Main-Wave Reward Redesign *(2026-05-06)*\n\n**Goal.** The forward_10d reward target is \"average 10-day forward log-return\". But what we actually care about is **realized return until a sensible exit signal fires**. Implement an MA5/MA10 death-cross exit with a 5-day hard cap.\n\n**Method.** New module `src/aurumq_rl/main_wave_labels.py` computes per-stock per-day `hold_return[t, j]` under signal-exit (`min(5d, MA5\u003cMA10 death cross)`). The env's reward function reads from this tensor instead of computing forward returns on-the-fly. The `valid_mask` is tightened to `entry_eligible \u0026 label_valid` so training reward and OOS evaluation filter on the same criterion.\n\n**Three runs (8h each on RTX 4070, short panel 2023-01..2025-06 train / 2025-07..2026-04 OOS).**\n\n- **22A**: seed=42, top_k=5, 300k steps\n- **22B**: seed=1, top_k=5, 300k steps (robustness check)\n- **22C**: seed=42, top_k=3 train, 200k steps (concentration check)\n\n**Key findings.**\n\n| Run | top_k | best step | hit_rate | win_rate | avg_hold | avg_dd | eval_score |\n|---|---:|---:|---:|---:|---:|---:|---:|\n| Phase 16a (V1 forward_10d, prod) | 3 | 224928 | 4.88 % | 36.9 % | +0.20 % | 3.00 % | +0.0490 |\n| Phase 21A (V2 forward_10d, REJECTED) | 3 | 149952 | 3.70 % | 34.5 % | −0.16 % | 4.71 % | −0.0596 |\n| **Phase 22A (V1 main_wave seed=42)** | 3 | 299904 | **5.89 %** | 41.1 % | +0.44 % | 3.68 % | +0.0419 |\n| Phase 22B (V1 main_wave seed=1) | 3 | 24992 | 6.06 %* | 40.2 % | +0.18 % | 3.79 % | +0.0168 |\n| **Phase 22C (V1 main_wave train_topk=3 → eval@5)** | 5 | 174944 | **6.16 %** | **44.0 %** | **+0.62 %** | 3.84 % | **+0.0505** |\n\n*22B's \"best\" is at step 24992 (≈ 1.5 PPO iterations); subsequent regress.\n\n**Three deltas vs Phase 16a baseline.**\n\n1. **Hit rate**: random base is 5.72 %. Forward_10d models all below random (−0.84 to −0.87 pp). main_wave models seed=42 series consistently **above random** (+0.14 to +0.44 pp).\n2. **Win rate**: +4 to +7 pp uniform lift across all three runs / all top_k variants.\n3. **Avg hold return**: 3× lift (16a +0.20 % → 22C +0.62 %).\n4. **Drawdown trade-off**: slightly higher `avg_max_drawdown` (3.65–3.95 % vs 16a's 3.00 %). The hold_return reward doesn't directly penalize drawdown — model is willing to ride larger in-hold drawdowns to capture larger gains. Net positive on eval_score, worth a Phase 23 fix.\n\n**Decision.** Phase 22C → production candidate. First time a model **clears the 5.72 % random baseline** on the main_wave criterion.\n\n### Phase 23 — Episode-target 清理 / Episode-Target Cleanup *(2026-05-06)*\n\n**Goal.** Clean up valid_mask edge cases and tighten the eval pipeline. Specifically:\n- last-week-of-data label_valid sometimes True because label window extends past panel end → forward fill gives +0 % return.\n- T+0 entry/exit on same day (entry day = exit day, hold_return = 0, treated as \"valid loss\") was double-counted.\n\n**Method.** Three patches in `main_wave_labels.py`: (a) strict label window check; (b) drop t==exit_t entries; (c) recompute `entry_eligible` to exclude T+0 (only buy after 9:30 of day t, exit at or after day t+1 close).\n\n**Result.**\n\n| Run | hit_rate | win_rate | avg_hold | eval_score |\n|---|---:|---:|---:|---:|\n| Phase 22C | 6.16 % | 44.0 % | +0.62 % | +0.0505 |\n| **Phase 23A (same config + cleanup)** | **6.42 %** | **44.6 %** | **+0.71 %** | **+0.0617** |\n\n**Production candidate locked.** Phase 23A is the first model that combines: (a) clears random base rate, (b) \u003e44 % win rate, (c) \u003e0.7 % avg hold return, (d) tight label semantics. Becomes \"23A baseline\" in subsequent phases.\n\n### Phase 24 / 25 — 技术因子改训 + 重要性权重 — 全部拒绝 / Tech-Factor Detour + Importance-Weighting REJECTED *(2026-05-07)*\n\n**Goal.** Add ~36 technical-analysis factors (MA / KDJ / MACD / Bollinger / amplitude / limit-up counts) on top of the 353-col baseline. Secondarily test per-factor importance-derived input weights.\n\n**Method.**\n- **Phase 24A**: compute tech factors at panel-load time inside the RL repo (KDJ/MACD computed from close-only because parquet had no OHLC). Train 300k seed=42 + `--add-technical-factors`.\n- **Phase 25A**: add IG-saliency × |T-1 z-score| sigmoid weights on top of 24A.\n- **Phase 25D**: weighting on the 353-col base WITHOUT tech, to isolate the weighting paradigm itself.\n\n**Results.**\n\n| Run | top-5 T-1 hit | lift vs random | Note |\n|---|---:|---:|---|\n| 23A baseline | 2.11 % | **2.38×** | the reference |\n| **Phase 24A (tech, 353+36=389 cols)** | **0.40 %** | **0.45×** | **below random 0.89 %** |\n| Phase 25A (24A + weighting) | 0.50 % | 0.56 % | VRAM-thrashed (96 % on RTX 4070), fps 172 → 4, killed at 60 % |\n| Phase 25D (weighting only, no tech) | 1.41 % | 1.59× | **−33 % vs 23A** |\n\n**Root causes** (post-mortem):\n\n1. **KDJ / amplitude approximated from close-only.** The parquet didn't carry OHLC at the time. `kdj_k = 100 * (close - min(close,9)) / (max(close,9) - min(close,9))` is a degenerate variant of true KDJ. Z-scoring this approximation produces strong artifacts on quiet stocks.\n2. **Binary event flags pollute LayerNorm gradients.** `ma_cross_5_10`, `golden_cross` are 0/1 indicators. After z-score, \"1\" days become ~3-5σ outliers; LayerNorm scales them down, but back-propagated gradient hits these few outlier samples and produces large updates that destabilize the policy.\n3. **Weighted top-30 factors saturate encoder capacity.** The per-stock encoder is 64-dim. Forcing 30 high-weight factors through a 64-dim bottleneck kills fine-grained ensemble structure that previously distributed signal across the full 353 cols.\n\n**Decision.** **Both directions rejected.** Architecture rule re-affirmed: **factor computation belongs in the upstream parquet pipeline, never at panel-load time inside the RL repo**. The team wrote `TECH_FACTOR_SPEC.md` (203 lines) and handed it back to the data team for proper OHLC-based tech factors. 23A remained production. Importance-weighting paradigm permanently dropped.\n\n### Phase 26A → G — cyq 修复 + 事件衰减编码 / cyq Fix + Event-Decay Encoding *(2026-05-07 ~ 08)*\n\n**Phase 26 is a 7-step recovery and breakthrough chain.** Track it carefully — this is where the project's current production model came from.\n\n#### Phase 26A: 加入 v1.0 cyq + 30 tech 列 / Add v1.0 cyq + 30 tech cols\n\n**Result.** 373-col training (343 base + 30 upstream tech_/cmf_/zt_) with new \"canonical\" cyq replacing legacy 88%-NaN cyq from `quotes_enriched`. **Regressed**: T-1 lift = 1.36× (vs 23A 2.38×).\n\n#### Phase 26B-baseline: 移除 30 tech 列, 保留 v1.0 cyq / Remove tech, keep cyq\n\n**Result.** 343-col, no tech. Still regressed: T-1 lift = 1.47×. **The regression is NOT from tech factors.**\n\n#### Phase 26 root-cause analysis: cyq backfill regime shift\n\nTraced to **`cyq_perf` backfill regime split**:\n- v1.0 cyq table only had **real Tushare data ≥ 2025-10-20**; everything before was backfilled with a different methodology.\n- Backfill null rate **0.61 %** vs real **26.53 %**.\n- `cyq_cost_distance` std: **0.197** (backfill, in training window) vs **0.066** (real, in OOS) — a **3× compression**.\n- Training was **100 % backfill**; OOS was **~63 % real**.\n- **Cross-section z-score does NOT equalize a mid-stream regime shift.** The model learned the synthetic-backfill scale; on real data the dispersion is 3× smaller, the model's \"this stock has high cyq\" signal collapses.\n\n**23A's accidental advantage.** 23A used the legacy 88%-NaN cyq, so the model had effectively learned to ignore cyq entirely. By contrast 26A/B's cleaner-but-distribution-shifted cyq actively misled the model.\n\n#### Phase 26C: 343-col + v1.2 cyq, wrong train window\n\n**v1.2 cyq fix from data team**: re-fetch via bulk Tushare API, all dates have consistent methodology. But: train end date was 2024-12-31 (vs 23A's 2025-06-30 — 6-month gap before OOS start).\n\n**Result.** T-1 lift 1.47× (same as 26B). Train-window mismatch swamps any cyq improvement.\n\n#### Phase 26C2: 353-col 23A-exact config + v1.2 cyq + correct train window ⭐\n\n**Result.** T-1 lift **2.61×** (2.31 % hit rate, best ckpt at **step 50k**). **+9.7 % over 23A's 2.38 %, AND converges 4× faster** (50k vs 200k). v1.2 cyq fix is validated. **Production candidate.**\n\n#### Phase 26D: 26C2 + 30 tech cols (clean panel)\n\n**Result.** T-1 lift **1.13×**. Adding 30 tech factors on the clean panel still hurts by −57 %. Confirms Phase 24's diagnosis: at 128→64→32 per-stock encoder capacity, the 30 raw tech cols dilute attention from strong alpha/gtja/mfp signals.\n\n#### Phase 26E/F/G: 事件衰减编码 / Event-Decay Encoding\n\n**Three new variants** at fresh 3-seed baseline:\n\n- **26E**: 26C2 (353 cols) + 2 continuous tech cols (`tech_boll_percent`, `cmf_120d_pct_amt`) = 355 cols\n- **26F**: 26E + 6 event-decay cols (`tech_evt_*` with τ=10d exp decay) = 361 cols\n- **26G**: 26F at bigger encoder (256→128, 256k params per-stock)\n\n3-seed median (seeds 42/43/44):\n\n| Run | factors | T-1 lift median | T-1 hit median | best ckpt |\n|---|---:|---:|---:|---|\n| 26C2 (clean panel sanity) | 353 | 1.70× | 1.50 % | step 50k |\n| 26E | 355 | 1.59× | 1.41 % | step 50k |\n| **26F (event-decay)** | **361** | **2.15×** | **1.90 %** | **step 50k** |\n| 26G (bigger encoder) | 361 | (abandoned) | — | — |\n\n**Best 26F seed**: seed=44 hit **2.72×** lift at step 50k (T-1 hit 2.41 %).\n\n**26G abandoned**: 256→128 encoder on RTX 4070 12 GB thrashed fps from 326 down to 4–55 with VRAM stranded by zombie contexts. 3 hours of attempts, no clean result. The capacity question is deferred to Linux / 16 GB-class hardware.\n\n#### Phase 26F-v3 / G-v3: clean-panel re-verification *(2026-05-08)*\n\nRe-run the comparison on the panel-v3 build (cyq v1.2 fully shipped, alpha/gtja sanitizer integrated upstream, mf_ `_log` variants emitted).\n\n| Run | T-1 lift median | range |\n|---|---:|---:|\n| 26C2-v3 sanity | 1.70× | 1.36–2.15× |\n| **26F-v3 ⭐ (PRODUCTION)** | **2.27×** | **2.04–2.38×** |\n| 26G-v3 | 1.82× | 1.59–2.05× |\n\n**Decision.** **26F-v3 = final RL production model.** 361 cols (353 base + 2 continuous tech + 6 event-decay) at 128→64 encoder. Bigger-encoder hypothesis closed (disproven on clean panel too — not a hardware artifact).\n\n**Phase 26 lesson summary.**\n\n1. The win came from **continuous event-decay encoding of binary signals** (the explicit fix to Phase 24's binary-flag failure), not from continuous TA on close prices.\n2. Mid-stream regime shifts (cyq v1.0 backfill vs real) defeat cross-section z-score; only the data team can fix it (v1.2 bulk-API recompute).\n3. Train-window matters as much as architecture. A 6-month gap between train-end and OOS-start is enough to wipe out any cyq-quality improvement.\n\n---\n\n## §6 监督学习对照赛道（paris 侧）\u003ca id=\"6-监督学习对照赛道paris-侧\"\u003e\u003c/a\u003e\n## §6 Supervised-Learning Companion Track \u003ca id=\"6-supervised-learning-companion-track\"\u003e\u003c/a\u003e\n\n**中文.** 与 RL 赛道并行，paris 侧（AurumQ 主仓）维护了一条 LightGBM/CatBoost/XGBoost 监督学习赛道作为对照。这条赛道**已经超越了 RL 赛道的实证表现**，但 RL 仍然作为多样性来源继续生产。下文是 SL 赛道关键节点的精简记录。\n\n**English.** In parallel with the RL track, paris (AurumQ main repo) maintains a LightGBM/CatBoost/XGBoost supervised-learning track as control. The SL track **has empirically outperformed the RL track**, but RL remains in production as a diversity source. Below is a compressed record of the SL track's key milestones.\n\n### 6.1 P0 — Wave Label 消融 / Wave Label Ablation *(2026-05-09)*\n\n**Goal.** Among four labeling methods (A v2_excess_adaptive, B trend-scanning, C triple-barrier, D directional-change), pick the production main-wave label.\n\n**Method.** Calibrate every method to train pos_rate ≈ 0.80 %, train LGBM (3 horizons each — t1, t3, e20) on the 26F-v3 348-col panel, score on test 2025-07..2025-12. Decision weights: 0.45·PR_AUC + 0.20·(1−ECE) + 0.15·top1 %_lift + 0.10·(1−ind_cv) + 0.10·(1−year_cv).\n\n**Stage 2 test PR-AUC at t3:**\n\n| Method | best_iter | PR_AUC | lift | Brier_ratio | ECE | top1 % | top5 % | daily_prec@5 |\n|---|---:|---:|---:|---:|---:|---:|---:|---:|\n| **A** | 191 | **0.1217** | **3.0×** | 0.972 | **0.010** | 3.65× | 4.21× | **0.200** |\n| B | 196 | 0.1052 | 2.4× | 0.977 | 0.013 | 4.16× | 3.35× | 0.156 |\n| C | 162 | 0.1195 | 2.5× | 0.968 | 0.015 | 4.25× | 3.70× | 0.203 |\n| D | 8 | 0.0961 | 2.5× | 0.971 | 0.005 | 2.31× | 3.25× | 0.114 |\n\nComposite-decision score: A = 0.836, C = 0.815 (within tiebreak band 0.03). C wins industry_cv tiebreak (0.40 vs 0.51); A wins on operational clarity + 15d vs 5d median event duration.\n\n**Null tests** both **PASS**: label-shuffle PR_AUC = 0.04021 (0.989× base), date-shuffle 0.06060 (1.491× — borderline, threshold 1.5×). Real model lift 3.0× / shuffle lift 1.49× → **2.0× fresh-signal ratio**.\n\n**Decision.** P0 = **Method A at horizon t3**, threshold τ_A = 1.2327, locked in `src/aurumq/labeling/p0_chosen.py`. C kept as drop-in fallback. P0 cleared for P1 production deployment.\n\n### 6.2 P2 — Production Hardening *(2026-05-09)*\n\n**Goal.** Close P1's training gaps so the LGBM is production-grade, not an ablation prototype.\n\n**Pieces.**\n\n- **Stage 0**: Daily panel rebuild. Split monolithic `feature_panel_v3_344.parquet` (3.65 GB) into 804 per-day shards (`year=YYYY/date=YYYY-MM-DD.parquet`). Celery beat @18:35 between phase20 rebuild (18:30) and wave_scores (18:45). Schema-hash assert against P0 lock `5e71e158e331`.\n- **Stage 1**: Walk-forward on single anchor 2025-07, **5 seeds × 3 horizons** (t1, t3, e20) = 15 LGBM trainings. LGBM params locked from P0 (lr=0.02, num_leaves=63, n_est=1500, early_stop=80), per-seed isotonic, mean ensemble.\n- **Stage 2**: composite_mean(A, C) finishing — **REJECTED** (PR_AUC = 0.0433 vs A_t3's 0.1217, Δ = −0.0785). Composite labels mix A∪C event positions that don't overlap, raising noise.\n- **Stage 3**: PPO residual — SKIPPED, deferred to P3 4070 work.\n- **Stage 4**: Alembic 052, v1/v2 dual-write, `drift_check @19:00`.\n\n| Horizon | test_pos_rate | PR_AUC | lift | ECE | top1 % | daily@5 | per-seed std |\n|---|---:|---:|---:|---:|---:|---:|---:|\n| **t1** | 0.0135 | **0.0721** | **5.34×** | 0.0024 | **9.28×** | 0.103 | 0.002 |\n| **t3** | 0.0407 | 0.1224 | 3.01× | 0.0100 | 3.33× | **0.209** | 0.001 |\n| **e20** | 0.2650 | 0.4136 | 1.56× | 0.0477 | 1.96× | 0.548 | 0.002 |\n\n**Decision.** Promote `wave_t3_lgbm_v1` (P0 single seed) → `wave_t3_lgbm_v2.anchor=2025-07.ensemble` (5 seeds) under tiered \"realistic gate\": t3 + t1 ship, e20 reference-only.\n\n### 6.3 Multi-Path SL Exploration: Path 1 / 2 / 4 / 5 / 6 *(2026-05-10)*\n\n**Diversity exploration.** 4 new bundles + 1 meta-stack on the same H1/H2 windows.\n\n| Path | Input panel | Preprocessing | n_features | Model |\n|---|---|---|---:|---|\n| Path 1 | feature_panel_v3_344 | **NO rank-z** (raw) | 345 | LightGBM β-regression |\n| Path 2 | feature_panel_clean | rank-z (same as Path 4) | 345 | **CatBoost + XGBoost mix** |\n| Path 4 | feature_panel_clean | rank-z | 345 | LightGBM (prod) |\n| Path 6 | feature_panel_clean_pruned | rank-z + drop 119 SHAP-zero | 226 | LightGBM (Bayesian opt) |\n| Path 5 | meta-LGB over (Path1, Path4, Path2) | + 11 regime + 9 interactions | 23 meta | LightGBM meta |\n\n**Short-panel scoreboard** (H1 = 2025-07..09, H2 = 2025-10..12):\n\n| Path | H1 cal primary | H2 cal primary | T1_hit H1 |\n|---|---:|---:|---:|\n| Path 1 | +0.028030 | +0.030497 | 54.4 % |\n| **Path 4** | **+0.028483** | **+0.030577** | 54.5 % |\n| Path 2 | +0.027932 | +0.030658 | 54.0 % |\n| Path 6 | +0.028265 | +0.030739 | 54.5 % |\n| **Path 5 (meta)** | +0.028372 | +0.030245 | **55.8 %** |\n| Path D (long, REJECTED as standalone) | +0.028358 | +0.030738 | — |\n| Path 3 (TabNet, **REJECTED**) | +0.019417 | +0.020 | — |\n\n**Path 3 TabNet's rejection.** H1 +0.019 vs Path 1 +0.028 (−30-40 %); 47 min train vs 50 s for LightGBM (~55× slower); killed 7/8 grid combos. Also drove T1_hit DOWN by ~1 pp when added to meta — excluded from Path 5 stacking. **TabNet is not the answer for this problem under this data scale.**\n\n**Cross-path ensemble**: rank-mean across Path 1+4+2 gives H1 +0.028407 vs best single +0.028483 — Δ = −0.0001. **Three GBDT-family paths are highly correlated; ensemble diversity is exhausted.**\n\n### 6.4 Long-Panel Sweep: The Climax *(2026-05-10 ~ 11)*\n\n**Goal.** Use the 7-year long panel (2018-01-02..2024-12-04 train, same H1/H2 as short) to retrain Path 1/2/4/5. Test the rank-z hypothesis.\n\n**Combined headline:**\n\n| Path | short H1 | long H1 | Δ H1 | short H2 | long H2 | Δ H2 |\n|---|---:|---:|---:|---:|---:|---:|\n| Path 1 (raw LGB) | +0.028030 | +0.028626 | **+5.97 bps** | +0.030497 | +0.031089 | +5.92 bps |\n| Path 4 (rank-z LGB) | +0.028483 | +0.028358 | **−1.25 bps** | +0.030577 | +0.030738 | +1.61 bps |\n| Path 2 (CB+XGB) | +0.027932 | +0.028343 | +4.11 bps | +0.030658 | +0.030857 | +2.00 bps |\n| Path 5 (regime stack) | +0.028372 | +0.028817 | +4.46 bps | +0.030245 | +0.031131 | **+8.86 bps** |\n\n**Finding 1 — rank-z kills long-panel info.** Experiment B isolates the variable: Path 4 hyperparams + the same raw long panel = **numerically identical to Path 1 long** (H1 = 0.028626). The per-day cross-sectional rank operation erases cross-year factor amplitude information — fine for 2y where amplitude variation is small, catastrophic for 7y where most of the new info lives in cross-year amplitude regimes.\n\n**Finding 2 — 5-year is the plateau** (Experiment A):\n\n| Train window | n years | H1 mean | H2 mean |\n|---|---:|---:|---:|\n| 2023-2024 (2y) | 2 | +0.027928 | +0.030285 |\n| 2022-2024 (3y) | 3 | +0.028182 | +0.030880 |\n| **2020-2024 (5y)** | 5 | **+0.028529** | **+0.030953** |\n| 2018-2024 (7y) | 7 | +0.028533 | +0.030983 (plateau) |\n\nStrictly monotonic up to 5y then flat. **2018-2019 contribute nothing.** Implications: if retraining, use 2020-2024, save 30 % compute.\n\n**Finding 3 — Strategy D compounds.** Top-50 score-weighted sizing:\n\n| Path | + Strategy D H1 mean_y | + Strategy D H2 mean_y |\n|---|---:|---:|\n| Path 4 short + Strategy D (current prod v2) | +0.0308 | +0.0333 |\n| **Path 1 long + Strategy D** | **+0.0315** | **+0.0343** |\n| Δ | **+7 bps** | **+10 bps** |\n\nStrategy D's +8 % concentration effect **stacks additively** on a stronger base — bigger absolute gains, not just same percentage of smaller pie.\n\n**Three production candidates:**\n\n| Candidate | H1 vs prod | H2 vs prod | + Strategy D vs prod v2 | Ops cost |\n|---|---:|---:|---:|---|\n| **C. Hybrid** (Path 1 long + Path 4 short 50/50) | +2.7 bps | +4.5 bps | +5-8 bps | **simplest** — 2 inferences + average, no new model |\n| **A. Path 1 long** (raw, 2018-2024) | +1.4 bps | +5.1 bps | **+7-10 bps** | medium — 1 new bundle |\n| **B. Path 5 long stacking** | +3.3 bps | +5.5 bps | strongest | high — 3 base + meta + 11 regime + 9 interactions |\n\n**All three were shipped to production on 2026-05-11.** Hybrid + path1_long went live first (commit `78d71ce`); path5_long followed 30 minutes later after ledashi shipped the missing path4_long + path2_long base bundles (commit `0ab6a55`). **Today path5_long is the H1-leading model** (+0.02882) and is one signal-A/B-window from being promoted to `is_recommended=True` over Path 4 short.\n\n---\n\n## §7 实证发现 — 六条改变方向的结论 \u003ca id=\"7-实证发现--六条改变方向的结论\"\u003e\u003c/a\u003e\n## §7 Empirical Findings — Six Conclusions That Changed Direction \u003ca id=\"7-empirical-findings\"\u003e\u003c/a\u003e\n\n### Finding 1 — Rank-z 会销毁长 panel 跨年幅度信号 / Rank-z Destroys Long-Panel Cross-Year Amplitude\n\n**中文.** 截面 z-score（rank-z）在每个 trade_date 内把当天因子重排到 [-σ, +σ] 区间。短 panel（2 年）训练时这没问题：所有因子的截面分布相似。长 panel（7 年）训练时灾难：2020 年和 2024 年的市场结构完全不同（科创板规模、北向资金占比、机构持仓比例都翻了几倍），而 rank-z 把跨年的「因子绝对幅度」信息全部抹掉，模型只能学到「相对排名」，反而损失 5-6 bps 的 H1 mean_y。\n\n**English.** Cross-section z-score (rank-z) re-ranks every factor's per-day distribution to [-σ, +σ]. Fine for short panels (2y) where cross-section distributions are similar. Catastrophic for long panels (7y) where market structure shifts dramatically year-over-year (STAR board size, Northbound share, institutional holdings all multi'd over 7y). Rank-z erases the cross-year factor amplitude information; the model only learns relative ranks, losing 5-6 bps H1 mean_y.\n\n**Action.** For long-panel training, **use raw features**. Verified: Path 1 long (raw, 7y) \u003e Path 4 long (rank-z, 7y) by 5.97 bps H1 with identical hyperparams.\n\n### Finding 2 — 5 年训练窗口是 plateau / Five Years is the Plateau\n\n**中文.** 2018-2019 数据零边际贡献。原因：A 股市场在 2019 年下半年到 2020 年初经历了科创板开板、注册制改革、外资准入扩大 —— 实质上是个 regime change。把 2018-2019 数据塞进训练集相当于让模型同时学两个市场结构，跨年泛化变差。\n\n**English.** 2018-2019 contributes nothing. Reason: A-shares went through a regime change in late 2019 / early 2020 (STAR opening, registration-based IPO reform, foreign-access expansion). Training on 2018-2019 forces the model to learn two market structures simultaneously, hurting cross-year generalization.\n\n**Action.** Default train window is **2020-2024 (5y)**. Saves 30 % compute at same precision.\n\n### Finding 3 — Strategy D 与任何 base 都正向叠加 / Strategy D Compounds Additively\n\n**中文.** Strategy D = top-K 仓位按校准分数加权（`weight = max(score, 0) / sum`），而不是等权。在 Path 4 short 上 +8% mean_y；在 Path 1 long 上 +10% mean_y。**新指标提升的绝对值 ≈ 旧 base 提升的绝对值**，不是百分比。换 base 越强，Strategy D 累加越多。\n\n**English.** Strategy D = top-K score-weighted position sizing (`weight = max(score, 0) / sum`), not equal-weight. Adds +8 % mean_y on Path 4 short, +10 % on Path 1 long. **Absolute improvement is the same whether base is weak or strong** — stronger base means larger combined gain. Always apply Strategy D.\n\n### Finding 4 — 二值事件标志必须 exp-decay 编码 / Binary Event Flags Must Be Exp-Decay Encoded\n\n**中文.** 二值事件（MA 金叉、KDJ 死叉、突破 3σ）原始 0/1 进 LayerNorm 是 −33% 准确率的元凶（Phase 24A）。正确做法：把二值序列做 `exp-decay τ=10d` 转换，`evt_decay(t) = Σ_{tau ≤ 10d} 1[event in last 10d] · exp(-(t-tau)/10)`，得到 0..1 连续衰减值。Phase 26F 用这个修复把 T-1 lift 从 1.13× 拉回 **2.27×**。\n\n**English.** Raw binary event flags (MA cross, KDJ death-cross, 3σ breakout) directly into LayerNorm cause a −33 % accuracy regression (Phase 24A). Correct encoding: exp-decay `τ=10d`. Phase 26F applied this fix and brought T-1 lift back from 1.13× to **2.27×**.\n\n### Finding 5 — Permutation importance 是 conditional, 不是 causal / Permutation Importance Is Conditional, Not Causal\n\n**中文.** Phase 16 跑出 cyq 和 inst 都是「robust 负向」（permutation importance 显著负）。Phase 17A 把它们 drop 掉重训 → adj Sharpe 暴跌 −0.732。原因：permutation importance 度量的是**当前训好的策略有多依赖某个特征**，不是**这个特征是否对预测有因果贡献**。要测因果必须**重训**而不是**重 permutation**。\n\n**English.** Phase 16 found cyq + inst both had \"robust negative\" permutation importance. Phase 17A dropped them and retrained → adj Sharpe collapsed −0.732. Reason: permutation importance measures **how much the trained policy depends on a feature**, not **whether the feature is causally helpful for prediction**. To test causality you must **retrain** after the drop, not just re-permute.\n\n### Finding 6 — 大多数 inf 是上游数据创伤, 不是公式 bug / Most Infs Are Upstream Data Wounds, Not Formula Bugs\n\n**中文.** Phase 26 跑 inf-root-cause 审计时，发现 19 个 inf-producing 因子的根因不是公式问题：\n\n1. **adj_factor 损坏**：5 个日期 + 2026 全年 16.5 万 NULL → `compute_alpha101.adj.fill_null(1.0)` 把不复权价拼到复权价上，制造假 regime shift。\n2. **一字板 high=low=close**：vwap-close、high-low 都 = 0，公式 div-zero。\n3. **rolling 相关性 window=2**：std=0 触发 NaN/inf。\n4. **rank^delta 当 rank=0 且 delta 极端**：`0^-large = inf`，gtja_017 实测 max 达 1.4×10³⁰⁸（fp64 上限）。\n\n**英文 lesson**: sanitizer（inf→NaN + clip ±1e6）只是兜底，真正的修复在上游数据管线。\n\n**English.** A Phase 26 inf-root-cause audit traced 19 inf-producing factors not to formula bugs but to four upstream data wounds: (1) adj_factor corruption (5 dates + all of 2026 = 165k NULL → `compute_alpha101.adj.fill_null(1.0)` stitched unadjusted prices to adjusted prices, creating fake regime shifts); (2) one-字板 days (high=low=close → vwap-close and high-low both zero, divide-by-zero); (3) rolling correlation window=2 producing std=0; (4) `rank^delta` with rank=0 and extreme delta blowing past fp64 (gtja_017 hit max 1.4×10³⁰⁸). **Lesson**: sanitizer (inf→NaN + clip ±1e6) is only a backstop; the real fix is upstream data.\n\n---\n\n## §8 生产流水线 — 每日 18:30-19:00 评分预算 \u003ca id=\"8-生产流水线--每日-1830-1900-评分预算\"\u003e\u003c/a\u003e\n## §8 Production Pipeline — Daily 18:30-19:00 Scoring Budget \u003ca id=\"8-production-pipeline\"\u003e\u003c/a\u003e\n\n**中文.** 当前生产环境（AurumQ 主仓侧）的 Celery Beat 调度（工作日）：\n\n```\n18:30  phase20_rebuild_panels_daily      重建 short combined panel + 上传 OSS\n18:35  rebuild_feature_panel_daily       当日 shard 写入\n18:45  generate_wave_scores_daily        v1/v2 集成（legacy 兼容）\n18:50  generate_wave_scores_path4_daily  Path 4 + Strategy D（当前推荐）\n18:51  path1 评分\n18:52  path2 评分\n18:53  path6 评分\n18:54  path1_long 评分                   ★ 长 panel raw\n18:55  path4_long 评分                   ★ path5_long base\n18:56  path5 评分                        meta on path1+path4+path2\n18:57  path2_long 评分                   ★ path5_long base\n18:58  hybrid 评分                       ★ path1_long + path4 50/50\n18:59  path5_long 评分                   ★ NEW BEST regime stack on long bases\n19:00  wave_drift_check                  漂移监控\n```\n\n**关键点**：\n- **10 个 model_version 并存**。任何模型上线，路径：`runner.py:PATH_CONFIG` 加一条 + `celery_beat.py:BEAT_SCHEDULE` 加一行 + `wave","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fyupoet%2Faurumq-rl","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fyupoet%2Faurumq-rl","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fyupoet%2Faurumq-rl/lists"}