majidraza1228
Community@majidraza1228
My impressive career of more than 15 years in software solutions architecture has been hallmarked by a genuine passion for developing and launching new products
Agent Skills by majidraza1228
Showing 9 vetted skills indexed across 1 GitHub repositories.
eval-coding-agent
Identify and quantify failure modes in coding agent outputs.
error-analysis
Analyze LLM pipeline traces to identify and categorize failure modes.
build-review-interface
Render LLM trace data in a browser for Pass/Fail labeling and JSONL export.
write-judge-prompt
Design binary Pass/Fail LLM-as-Judge prompts for defined failure modes.
eval-audit
Audit LLM eval pipelines and generate prioritized findings reports.
evaluate-rag
Evaluate RAG pipelines by measuring retrieval and generation quality separately.
validate-evaluator
Calibrate an LLM judge against human labels using train/dev/test splits.
eval-tool-use
Evaluate LLM agent tool selection, argument validation, and sequencing.
generate-synthetic-data
Generate synthetic traces with dimension-based tuples for LLM pipeline evaluation.