SIMULATION STANDARD · V0.1

Data, model and experiment requirements

Use these requirements before adding a dataset, registering a Python model or sharing an experiment. They keep results reproducible, comparable and safe to reuse when more members join the Simulation Lab.

繁體中文版 →

1. Dataset requirements

Accepted now: CSV, XLSX and XLS. Database, API, warehouse and enterprise-system connectors can be added later without changing the dataset contract.

Required itemWhat to provide
Dataset ID and versionA stable identifier and a new version whenever source data changes.
Source and observation windowWhere the data came from, start/end dates and timezone.
SchemaColumn name, business meaning, type, unit and role.
Keys and timePrimary/entity key and timestamp column when applicable.
TargetTarget column for supervised models, or clearly state that no target is used.
Missing-value ruleWhat blank, null or unknown means for each important field.
Quality statusDraft, Checked, Approved or Rejected.
ProvenanceOriginal file hash or source-system version, plus owner and notes.

Do not upload secrets, passwords or personal identifiers into the Lab. PII, confidential production data and regulated data require a separate approval and storage design before broader use.

2. Field dictionary

Every model-relevant column should have a short dictionary entry so a later user does not need to guess what a field means.

FieldRequired description
Column nameExact name in the file or database.
Business meaningPlain-language description of what the field represents.
Data typeString, integer, float, boolean, category, datetime or text.
RoleIdentifier, timestamp, feature, target, group, weight or ignored.
Unit / rangeUnit, allowed values or expected range where applicable.
Missing ruleDrop, impute, preserve as unknown or another documented rule.

3. Python model requirements

Model code is version-controlled and deployed through Codex / Git workflows. Ordinary browser users register an approved model version; they do not upload arbitrary Python for execution.

Required itemWhat to record
Model ID / versionStable model ID and semantic or otherwise explicit version.
Task typeClassification, regression, forecasting, anomaly detection, optimization, simulation, RL or other.
FrameworkRule-based, scikit-learn, CatBoost, TensorFlow/Keras, PyTorch, custom Python, etc.
Python environmentPython version plus requirements/lockfile or container environment.
Entry pointThe callable used by the runner, for example models/severity.py:run.
Git referenceRepository and exact commit used for the run.
Input / output contractExpected input schema and output structure.
Parameters / seedDefault parameters and random-seed policy.
MetricsTask-appropriate metrics and any business KPI used for interpretation.
LimitationsKnown exclusions, drift risks, data constraints and conditions where the model should not be used.
Approval stateDraft → Tested → Validated → Member-ready → Archived.

4. Experiment run requirements

Every saved run must point to exact versions. A result is not considered reproducible if any one of these references is missing.

Dataset version + Model version + Git commit + Cycle version + Scenario + Parameters + Seed + Environment + Timestamp + Output + Metrics + Notes

Past runs are never overwritten. A rerun or reassessment creates a new run ID.

5. Flexible decision cycle

The default Lab cycle is intentionally modular:

Question → Data → Prepare → Model → Scenario → Evaluate → Compare → Decide → Action → Review

Users may drag, add, remove, duplicate or rename visible steps. The audit record remains independent of the visual layout, so the system still keeps the data, model and run versions needed to reproduce the experiment.

6. What can be shared with other members?

StatusMeaningWho should use it
DraftRegistered but not checked.Lab owner / contributor only.
TestedTechnical execution completed.Lab team only.
ValidatedEvidence, metrics and limitations reviewed.Reviewer-approved use.
Member-readyApproved for broader controlled member use.Eligible members.
ArchivedRetained for reproducibility but no longer recommended.History / audit only.

Current private beta access is restricted. Future member roles can follow the same standard without changing the underlying dataset, model or experiment records.

7. Future member roles and version rules

RolePlanned responsibility
Member ViewerUse only datasets, models and cycles that have been marked Member-ready.
Lab ContributorAdd datasets, field dictionaries, cycles and draft model registrations; cannot approve promotion.
ReviewerReview validation evidence, metrics, limitations and reproducibility before promotion.
Lab OwnerManage access, approve promotion, archive assets and control the Lab configuration.

Current beta: only Lab Owner access is enabled. The additional roles are defined now so later users can be added without redesigning the data model.

Version rules

Do not silently edit an approved dataset, model or cycle. Any material change creates a new version. Every experiment run keeps immutable references to the exact dataset version, model version, Git commit and cycle version used.