SIMULATION STANDARD · V0.1
Data, model and experiment requirements
Use these requirements before adding a dataset, registering a Python model or sharing an experiment. They keep results reproducible, comparable and safe to reuse when more members join the Simulation Lab.
1. Dataset requirements
Accepted now: CSV, XLSX and XLS. Database, API, warehouse and enterprise-system connectors can be added later without changing the dataset contract.
| Required item | What to provide |
|---|---|
| Dataset ID and version | A stable identifier and a new version whenever source data changes. |
| Source and observation window | Where the data came from, start/end dates and timezone. |
| Schema | Column name, business meaning, type, unit and role. |
| Keys and time | Primary/entity key and timestamp column when applicable. |
| Target | Target column for supervised models, or clearly state that no target is used. |
| Missing-value rule | What blank, null or unknown means for each important field. |
| Quality status | Draft, Checked, Approved or Rejected. |
| Provenance | Original file hash or source-system version, plus owner and notes. |
Do not upload secrets, passwords or personal identifiers into the Lab. PII, confidential production data and regulated data require a separate approval and storage design before broader use.
2. Field dictionary
Every model-relevant column should have a short dictionary entry so a later user does not need to guess what a field means.
| Field | Required description |
|---|---|
| Column name | Exact name in the file or database. |
| Business meaning | Plain-language description of what the field represents. |
| Data type | String, integer, float, boolean, category, datetime or text. |
| Role | Identifier, timestamp, feature, target, group, weight or ignored. |
| Unit / range | Unit, allowed values or expected range where applicable. |
| Missing rule | Drop, impute, preserve as unknown or another documented rule. |
3. Python model requirements
Model code is version-controlled and deployed through Codex / Git workflows. Ordinary browser users register an approved model version; they do not upload arbitrary Python for execution.
| Required item | What to record |
|---|---|
| Model ID / version | Stable model ID and semantic or otherwise explicit version. |
| Task type | Classification, regression, forecasting, anomaly detection, optimization, simulation, RL or other. |
| Framework | Rule-based, scikit-learn, CatBoost, TensorFlow/Keras, PyTorch, custom Python, etc. |
| Python environment | Python version plus requirements/lockfile or container environment. |
| Entry point | The callable used by the runner, for example models/severity.py:run. |
| Git reference | Repository and exact commit used for the run. |
| Input / output contract | Expected input schema and output structure. |
| Parameters / seed | Default parameters and random-seed policy. |
| Metrics | Task-appropriate metrics and any business KPI used for interpretation. |
| Limitations | Known exclusions, drift risks, data constraints and conditions where the model should not be used. |
| Approval state | Draft → Tested → Validated → Member-ready → Archived. |
4. Experiment run requirements
Every saved run must point to exact versions. A result is not considered reproducible if any one of these references is missing.
Dataset version + Model version + Git commit + Cycle version + Scenario + Parameters + Seed + Environment + Timestamp + Output + Metrics + Notes
Past runs are never overwritten. A rerun or reassessment creates a new run ID.
5. Flexible decision cycle
The default Lab cycle is intentionally modular:
Question → Data → Prepare → Model → Scenario → Evaluate → Compare → Decide → Action → Review
Users may drag, add, remove, duplicate or rename visible steps. The audit record remains independent of the visual layout, so the system still keeps the data, model and run versions needed to reproduce the experiment.
6. What can be shared with other members?
| Status | Meaning | Who should use it |
|---|---|---|
| Draft | Registered but not checked. | Lab owner / contributor only. |
| Tested | Technical execution completed. | Lab team only. |
| Validated | Evidence, metrics and limitations reviewed. | Reviewer-approved use. |
| Member-ready | Approved for broader controlled member use. | Eligible members. |
| Archived | Retained for reproducibility but no longer recommended. | History / audit only. |
Current private beta access is restricted. Future member roles can follow the same standard without changing the underlying dataset, model or experiment records.
7. Future member roles and version rules
| Role | Planned responsibility |
|---|---|
| Member Viewer | Use only datasets, models and cycles that have been marked Member-ready. |
| Lab Contributor | Add datasets, field dictionaries, cycles and draft model registrations; cannot approve promotion. |
| Reviewer | Review validation evidence, metrics, limitations and reproducibility before promotion. |
| Lab Owner | Manage access, approve promotion, archive assets and control the Lab configuration. |
Current beta: only Lab Owner access is enabled. The additional roles are defined now so later users can be added without redesigning the data model.
Version rules
Do not silently edit an approved dataset, model or cycle. Any material change creates a new version. Every experiment run keeps immutable references to the exact dataset version, model version, Git commit and cycle version used.