1

Upload

Drop your raw files into the refinery. We currently support CSV, JSON, JSONL, and TXT (up to 50 KB on the free tier). Parquet support is on the roadmap.

You can also give optional natural-language instructions — for example: “Extract sentiment, standardize all dates to ISO-8601, and remove email addresses.”

Tip: The cleaner your source structure, the more precise the AI refinement becomes.
2

AI Refine

Choose your AI provider and mode. SchemaForge currently supports:

  • Google Gemini — strong general reasoning and structured output
  • Groq — extremely fast inference
  • Cerebras — high-throughput processing
  • SambaNova — enterprise-grade performance

Modes range from a full end-to-end pipeline to lighter cleaning or schema-only inference. The model analyzes structure, types, missing values, inconsistencies, and applies your instructions.

3

Validate

After refinement, SchemaForge runs automated quality checks:

Schema Inference
Column types, nullability, constraints
Consistency
Format alignment & value ranges
Completeness
Missing data detection
Quality Score
Overall readiness metric

You see a clear quality score and can inspect what changed before publishing.

4

Publish

Once validated, your dataset is ready for the real world:

  • Download as clean CSV / JSON / JSONL
  • Push directly to Hugging Face datasets
  • Keep private or share with the community

Every job is tracked so you can re-run, compare versions, or refine further with new instructions.

Try it now

After refinement — trust layer

Every completed dataset receives a Dataset DNA fingerprint, optional synthetic augmentation, a public telemetry snapshot, and an auto-generated README card. You can verify integrity anytime and publish to Hugging Face when ready.

View live telemetry Browse datasets