Kubesense

Playground

Use the playground to test a prompt with supported providers and models before changing production traffic. The playground can work with text or chat-style messages, structured output, and tool definitions. A trace can be opened as a starting point and its Langfuse-style input is converted into playground prompt messages.

LLM Prompt Hub playground

The playground lets you:

  • Choose a model from the selected application's configured connections.
  • Edit system, user, assistant, and tool messages.
  • Substitute prompt variables and placeholder message values.
  • Configure temperature, maximum tokens, and top-p sampling when supported.
  • Add tool definitions and test tool-call behavior.
  • Request structured output using a JSON schema.
  • Add another prompt instance to compare models or prompt variants side by side.
  • Reset the form or run the current prompt.

Playground example

To compare two models for a support response:

  1. Add a system message defining the support policy.
  2. Add a user message with {{question}} and provide a sample question.
  3. Select the first model and run the prompt.
  4. Add a second prompt instance, keep the same messages, and select another model.
  5. Compare output quality, latency, token usage, and cost.
  6. Save the better variant as a new prompt version and evaluate it on a representative dataset.

For structured output, request a response such as:

{
	"category": "billing",
	"priority": "high",
	"summary": "The customer was charged twice."
}

Define the corresponding schema in the playground so invalid responses are easier to identify before the prompt reaches production.

Playground limitations

The playground is a development aid. Its output can differ from production because of provider availability, model version, temperature, retrieved context, tool results, and application code around the prompt. Validate a candidate prompt with a representative dataset before deploying it.

From a trace to a prompt test

Open a representative generation from Tracing, review the exact input and output, and send the input to the playground. Adjust the prompt or model, compare the result, and record the candidate in a dataset when it should be evaluated repeatedly.

Where to next

  • Prompts — save a tested variant as a new version
  • Experiments — run a version across a dataset instead of one input