tokens&
For enterprises
Submit a resource
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Home
  2. Agent Skills
  3. NeMo Data Designer
NVIDIAModels
SKILL.md
License identified

Agent Skill

NeMo Data Designer

Create declarative synthetic-data generation pipelines for model development, evaluation, and retrieval workflows.

synthetic-datanemodatasets
Install this skillView repository

Package facts

Sourced from the vendor's own repository.

Vendor
NVIDIA
Category
Models
License
Apache-2.0 / CC-BY-4.0
License review
License identified
Supported clients
Codex, Cursor, Claude Code, VS Code, GitHub Copilot, Gemini CLI
SKILL.md size
4,710 bytes

Source snapshot: 2026-09-04. This listing does not verify installation, security, vendor participation, or product use.

Raw SKILL.mdSource repositoryInstall the Tokens& Agent Pack

Skill specification

Declared by NVIDIA in the package front matter. Trigger conditions are what the coding agent matches on before it loads the skill.

NeMo Data Designer SKILL.md front matter fields
Skill namedata-designer
Trigger conditionsUse when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.
Argument hintdescribe the dataset you want to generate
Declared licenseApache-2.0

Install data-designer

In a terminal with Node.js, npm and Git, run the command for your agent. The Skills CLI installs the complete package directory, including referenced files within it. Review its install prompt, then start a new agent session. A skill package does not set up an MCP server connection.

Claude Code

.claude/skills/data-designer/SKILL.md

Project skills are committed with the repo. Use the user directory for a personal install across every project.

Project install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'claude-code'

Personal install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'claude-code' --global

Codex

.agents/skills/data-designer/SKILL.md

Codex reads `.agents/skills/` as its primary location, which is also the cross-platform default other clients honour.

Project install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'codex'

Personal install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'codex' --global

Cursor

.agents/skills/data-designer/SKILL.md

Cursor also loads `.agents/skills/`, `.claude/skills/`, and `.codex/skills/`, so one committed copy can serve several clients.

Project install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'cursor'

Personal install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'cursor' --global

Gemini CLI

.agents/skills/data-designer/SKILL.md

Gemini CLI reads `.agents/skills/` first when both directories exist.

Project install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'gemini-cli'

Personal install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'gemini-cli' --global

GitHub Copilot

.agents/skills/data-designer/SKILL.md

The Skills CLI uses the shared `.agents/skills/` directory for Copilot project installs.

Project install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'github-copilot'

Personal install

npx skills add 'https://github.com/NVIDIA/skills/tree/main/skills/data-designer' --skill 'data-designer' --agent 'github-copilot' --global

SKILL.md

View raw source

Source snapshot fetched 2026-09-04 from github.com/NVIDIA/skills/tree/main/skills/data-designer. The install command fetches the upstream package, which may have changed since this snapshot.

Before You Start

Do not explore the workspace first. The workflow's Learn step gives you everything you need.

Goal

Build a synthetic dataset using the Data Designer library that matches this description:

$ARGUMENTS

Workflow

Use Autopilot mode if the user implies they don't want to answer questions — e.g., they say something like "be opinionated", "you decide", "make reasonable assumptions", "just build it", "surprise me", etc. Otherwise, use Interactive mode (default).

Read only the workflow file that matches the selected mode, then follow it:

  • Interactive → read workflows/interactive.md
  • Autopilot → read workflows/autopilot.md

Rules

  • Keep all columns in the output by default. The only exceptions for dropping a column are: (1) the user explicitly asks, or (2) it is a helper column that exists solely to derive other columns (e.g., a sampled person object used to extract name, city, etc.). When in doubt, keep the column.
  • Do not suggest or ask about seed datasets. Only use one when the user explicitly provides seed data or asks to build from existing records. When using a seed, read references/seed-datasets.md.
  • When the dataset requires person data (names, demographics, addresses), read references/person-sampling.md.
  • If a dataset script that matches the dataset description already exists, ask the user whether to edit it or create a new one.

Usage Tips and Common Pitfalls

  • Sampler and validation columns need both a type and params. E.g., sampler_type="category" with params=dd.CategorySamplerParams(...).
  • Jinja2 templates in prompt, system_prompt, and expr fields: reference columns with {{ column_name }}, nested fields with {{ column_name.field }}.
  • `SamplerColumnConfig`: Takes params, not sampler_params.
  • LLM judge score access: LLMJudgeColumnConfig produces a nested dict where each score name maps to {reasoning: str, score: int}. To get the numeric score, use the .score attribute. For example, for a judge column named quality with a score named correctness, use {{ quality.correctness.score }}. Using {{ quality.correctness }} returns the full dict, not the numeric score.

Troubleshooting

  • `data-designer` CLI not found: Tell the user that data-designer is not installed in this environment (requires Python >= 3.10). Ask if they would like you to create a virtual environment and install it, or if they prefer to do it themselves. Do not install anything without the user's permission.
  • Network errors during preview: A sandbox environment may be blocking outbound requests. Ask the user for permission to retry the command with the sandbox disabled. Only as a last resort, if retrying outside the sandbox also fails, tell the user to run the command themselves.

Output Template

Write a Python file to the current directory with a load_config_builder() function returning a DataDesignerConfigBuilder. Name the file descriptively (e.g., customer_reviews.py). Use PEP 723 inline metadata for dependencies.

# /// script
# dependencies = [
#   "data-designer", # always required
#   "pydantic", # only if this script imports from pydantic
#   # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field


# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
    field_one: str = Field(description="...")
    field_two: int = Field(description="...")


# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
    required_columns=["col_a"],
    side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
    # add custom logic here that depends on "col_a" and update row in place
    row["name_in_custom_column_config"] = "custom value"
    row["extra_col"] = "extra value"
    return row


def load_config_builder() -> dd.DataDesignerConfigBuilder:
    config_builder = dd.DataDesignerConfigBuilder()

    # Seed dataset (only if the user explicitly mentions a seed dataset path)
    # config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))

    # config_builder.add_column(...)
    # config_builder.add_processor(...)

    return config_builder

Only include Pydantic models, custom generators, seed datasets, and extra dependencies when the task requires them.

Add the registry badge

Maintainers can link this listing from the skill's own README. Free, no account needed, and it points back at the rendered package for anyone browsing the repo.

Markdown

[![NeMo Data Designer on tokens&](https://tokensand.com/api/badges/skill/nvidia-data-designer)](https://tokensand.com/agent-skills/nvidia-data-designer)

HTML

<a href="https://tokensand.com/agent-skills/nvidia-data-designer" target="_blank" rel="noopener">
  <img src="https://tokensand.com/api/badges/skill/nvidia-data-designer" alt="NeMo Data Designer on tokens&" />
</a>

More NVIDIA Agent Skills

All Agent Skills

AI-Q Blueprint deployment

Install, run, validate, troubleshoot, and stop a local or self-hosted NVIDIA AI-Q Blueprint environment.

Agents

CUDA-Q onboarding guide

Install CUDA-Q, validate simulators and hardware targets, and build reproducible quantum applications.

Models

cuOpt installation

Select and verify a compatible cuOpt Python, C, or REST server installation for an NVIDIA GPU environment.

Models

cuOpt numerical optimization API

Solve linear, mixed-integer, and quadratic programs with the cuOpt Python API and result diagnostics.

Models

cuOpt optimization formulation

Translate business constraints and objectives into verifiable cuOpt mathematical programs before implementation.

Models

cuOpt routing API for Python

Build vehicle-routing and fleet-optimization models with constraints, objectives, and solution validation.

Models