di-agent-knowledge-engine-streamsets

Manage StreamSets Data Collector engine deployment and job operations for IBM watsonx.data.

3|Updated May 1, 2026
One-click install
npx skills add https://github.com/IBM/ibm-watsonx-data-integration-skills --skill di-agent-knowledge-engine-streamsets
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: di-agent-knowledge-engine-streamsets
Source: https://github.com/IBM/ibm-watsonx-data-integration-skills/tree/main/agent/skills/di-agent-knowledge-engine-streamsets
Command: npx skills add https://github.com/IBM/ibm-watsonx-data-integration-skills --skill di-agent-knowledge-engine-streamsets

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill removes the guesswork from deploying, configuring, and operating StreamSets Data Collector engines in IBM watsonx.data integration.

Core Features & Use Cases

  • Environment setup: Create and tune StreamSets environments with engine versions, stage libraries, external resources, VPC allocation, and resource thresholds.
  • Engine operations: Start engines with Docker or Podman, choose tunneling or direct communication, verify health, and inspect logs.
  • Job reliability: Run, stop, reset, and resume jobs while managing offsets, failover, high availability, and multi-engine capacity.
  • Use case: An operator can use this Skill to provision a production-ready engine fleet, diagnose startup issues, and keep continuous data flows running during maintenance or failures.

Quick Start

Ask for help creating or troubleshooting a StreamSets environment, then follow the generated steps to deploy engines and run jobs safely.

Frequently Asked Questions about di-agent-knowledge-engine-streamsets

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I start a StreamSets Data Collector engine using Docker or Podman?

To start a StreamSets Data Collector engine, you need to configure your container runtime, either Docker or Podman, and execute the version-aware startup commands provided for your specific engine deployment.

How does high availability failover work for StreamSets jobs in IBM watsonx.data?

High availability failover for StreamSets jobs keeps continuous data flows running during maintenance or failures by managing multi-engine capacity and automatically transferring job operations across your active engine fleet.

How do I configure StreamSets environment resource thresholds and VPC allocation?

You can configure StreamSets environments by defining engine versions, allocating VPC resources, and setting resource thresholds to ensure your data collector engines have the necessary capacity for reliable job operations.

What is the best way to manage offsets when resetting or resuming StreamSets jobs?

The best way to manage offsets when resuming StreamSets jobs is to use the engine's job reliability features to stop, reset, and run jobs while preserving offset management state for continuous data integration.

Why does my StreamSets engine startup fail during tunneling or direct communication setup?

StreamSets engine startup may fail during tunneling or direct communication setup if task credentials lack network access or if container runtime configurations are incorrect for your corporate infrastructure environment.

Do I need specific credentials and network access to run StreamSets engines in corporate infrastructure?

Yes, running StreamSets engines in corporate infrastructure requires task credentials, network access, and proper container runtime configuration to ensure engines can communicate and execute jobs reliably.