caveman

Compress assistant responses to reduce token usage while preserving technical accuracy.

2|Updated Mar 16, 2026
One-click install
npx skills add https://github.com/hamzaPixl/pixl-ai --skill caveman-hamzapixl
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: caveman
Source: https://github.com/hamzaPixl/pixl-ai/tree/main/packages/crew/skills/caveman
Command: npx skills add https://github.com/hamzaPixl/pixl-ai --skill caveman-hamzapixl

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Long, wordy assistant replies waste tokens, increase cost, and slow down developer workflows; caveman compresses language to deliver concise, accurate technical responses that cut token usage dramatically.

Core Features & Use Cases

  • Intensity Levels: Select lite, full, ultra, or wenyan variants to control compression and stylistic register.
  • Persistent Mode: Remains active across turns until explicitly turned off, reducing repeated prompts to be brief.
  • Safety-aware Compression: Automatically suspends terse style for security warnings, irreversible actions, multi-step clarifications, and code/commit outputs.
  • Use Cases: Save API tokens in long debugging sessions, get compact code review comments, produce terse documentation summaries, or generate classical Chinese terse outputs for stylistic needs.

Quick Start

Switch to caveman lite to enable terse, token-efficient responses.

Frequently Asked Questions about caveman

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I reduce token usage in LLM responses during long debugging sessions?

You can reduce token usage by applying a terse compression style to LLM outputs, which strips filler language while preserving technical substance. This persists across turns until toggled off, cutting costs in extended developer workflows.

Can I control the level of brevity in technical responses?

Yes, you can control brevity using selectable intensity levels: lite, full, ultra, and wenyan variants. These options adjust the compression and stylistic register of technical outputs to match your specific token-saving needs.

Does terse compression affect code review comments and documentation summaries?

Terse compression applies to code review comments and documentation summaries by delivering compact, precise wording. It preserves technical accuracy while substantially reducing word count and token consumption.

Are security warnings and destructive operations affected by token-efficient compression?

No, token-efficient compression automatically suspends its terse style for security warnings, irreversible actions, and multi-step clarifications. This safety-aware mechanism ensures critical alerts remain fully detailed and unabbreviated.

What is the wenyan variant for concise LLM outputs?

The wenyan variant is a selectable intensity level that generates classical Chinese terse outputs. It provides a stylistic compression option for users needing token-efficient responses in that specific linguistic register.

When should I not use token-saving compression for assistant outputs?

You should not use token-saving compression when generating code, commit outputs, or multi-step clarifications. The system automatically respects these safety exceptions to prevent ambiguity in destructive operations and technical artifacts.