# Competition Lab

PrizeAI向けの、表形式コンペの提出前検証・ローカル評価・実験記録を行うCLIです。Node.js標準ライブラリだけで動き、ネットワーク接続、Kaggle API、認証情報、提出機能は持ちません。

固定問題のデモではなく、ID列と任意の1列以上の目的変数を指定して、回帰・二値分類・多クラス分類のCSVを扱います。ただし、画像・音声・時系列の特徴量生成や学習は担当しません。あくまで結果CSVの検証と比較を再現可能にする道具です。

## 必要環境

- Node.js 18以上（追加パッケージ不要）

## 起動

```powershell
cd <competition-lab.js を保存したフォルダ>
node .\competition-lab.js --help
```

## 主な用途

### 1. 提出CSVの検証

重複ID、空の目的変数、列不足、サンプル提出とのスキーマ差異を提出前に止めます。

```powershell
node .\competition-lab.js validate `
  --input .\submission.csv `
  --sample .\sample_submission.csv `
  --id-column id `
  --target-column prediction `
  --require-numeric
```

複数の予測列も指定できます。

```powershell
node .\competition-lab.js validate --input .\submission.csv --id-column id `
  --target-column target_a --target-column target_b --require-numeric
```

### 2. 正解CSVによる評価

IDで照合してから評価します。デフォルトの`auto`判定は小さな数値ラベルを分類と解釈するため、曖昧な場合は`--task`を明示してください。

```powershell
# 回帰: MAE / RMSE / R2
node .\competition-lab.js evaluate --truth .\truth.csv --prediction .\candidate.csv `
  --id-column id --target-column target --task regression

# 二値確率: accuracy / macro F1 / log loss
node .\competition-lab.js evaluate --truth .\truth.csv --prediction .\candidate.csv `
  --id-column id --target-column target --task binary --positive-label 1

# ラベル分類: accuracy / macro F1
node .\competition-lab.js evaluate --truth .\truth.csv --prediction .\candidate.csv `
  --id-column id --target-column class --task multiclass
```

`--report path\to\report.json`を付けると評価結果JSONも保存できます。

### 3. 候補の比較

同じ正解CSVに対して複数候補を比較し、Markdownレポートを作ります。回帰はRMSE昇順、二値確率はlog loss昇順、それ以外の分類はaccuracy降順です。

```powershell
node .\competition-lab.js compare --truth .\truth.csv `
  --candidate baseline=.\baseline.csv `
  --candidate model_b=.\model_b.csv `
  --id-column id --target-column target --task regression `
  --report .\comparison.md
```

### 4. 実験記録

学習コマンド、データ版、任意の数値指標をローカルJSONLに追記し、Markdown一覧にします。

```powershell
node .\competition-lab.js record --ledger .\runs.jsonl --name catboost-v4 `
  --data-version 2026-10-07 --metric rmse=10.603 --note "local CV"
node .\competition-lab.js summary --ledger .\runs.jsonl --report .\runs.md
```

## テスト

```powershell
node --test .\tests\competition-lab.test.js
```

## 実CSVを安全に評価する流れ

`split` は、ID（または指定したグループ）をまたがせずに、再現可能な学習・検証CSVを作ります。入力の順番に依存せず、`--seed`で同じ分割を再現できます。グループをまたぐ重複はリークになり得るため、利用できる場合は必ず `--group-column` を指定してください。

```powershell
# 例: group単位で25%を検証用に分け、監査結果も保存する
node .\competition-lab.js split `
  --input .\data.csv --id-column id --group-column customer_id `
  --target-column target --feature-column age --feature-column amount `
  --valid-fraction 0.25 --seed 42 `
  --train-output .\work\train.csv --valid-output .\work\valid.csv `
  --report .\work\split-audit.json

# 明示的な監査。ID/グループ重複、目的変数の丸写し列をエラーにする。
node .\competition-lab.js audit `
  --train .\work\train.csv --valid .\work\valid.csv `
  --id-column id --group-column customer_id --target-column target `
  --feature-column age --feature-column amount --report .\work\audit.json
```

`audit` は全てのリークを証明できるものではありません。ID・グループ重複と、特徴量が目的変数を全行で丸写ししているケースをエラーにします。学習・検証間で完全一致する特徴量行は警告として記録します。時系列の未来情報、外部データ、業務知識上のリークは利用者が確認してください。

### 軽量ベースライン

`baseline` は回帰では学習CSVの平均、分類では最多クラスだけを出力する軽量な基準線です。特徴量を学習するモデルではありませんが、CSV形式・分割・評価経路が動くことを確かめるために使えます。

```powershell
# 回帰または分類で、検証CSV向けの予測CSVと実行記録を保存する
node .\competition-lab.js baseline `
  --train .\work\train.csv --test .\work\valid.csv `
  --output .\work\baseline.csv --id-column id --target-column target `
  --task regression --report .\work\baseline.json

node .\competition-lab.js evaluate `
  --truth .\work\valid.csv --prediction .\work\baseline.csv `
  --id-column id --target-column target --task regression `
  --report .\work\baseline-evaluation.json
```

同じコマンドで `--task binary` または `--task multiclass` を指定すれば分類にも使えます。実用モデルの出力CSVは `baseline` の代わりに `evaluate` / `compare` に渡してください。

CSV解析、重複IDの拒否、回帰、二値確率、候補比較、実験台帳をテストします。

## 境界

- Kaggle提出、外部API呼び出し、資格情報の読込・保存はしません。
- 競技規約、リーク、外部データ可否はこのツールで判定しません。利用者が各コンペ規約を確認してください。
- ローカル真値で評価できる場面の支援用です。公開LBやPrivate LBを代替・推定するものではありません。
