Skip to main content

profClaw API Testing Status

Last Updated: 2026-02-05 19:30 CST

Test Results Summary

Architecture Change: SDK-Managed Multi-Step (v3)

Executor refactored from manual while loop (maxSteps: 1 + manual message accumulation) to AI SDK native multi-step (generateText with stopWhen + onStepFinish). Tool chaining now handled entirely by the SDK.

Quick Tests (Core)

Agentic Multi-Step Tests

Full Suite (./run-all.sh)

Unit Tests

SDK Multi-Step Refactor (v3)

What Changed

  • Removed: Manual while loop, executeStep(), processToolCalls(), injectContext(), buildStepContext(), manual message accumulation, manual tool result formatting
  • Added: wrapToolsWithExecute(), onStepFinish callback, stopWhen: [stepCountIs(N), hasToolCall('complete_task')]
  • Result: ~150 lines deleted, ~40 lines added. SDK handles message format and result feeding internally.

Key Improvements

  • Tool chaining works natively: SDK feeds tool results back as properly formatted messages
  • Proper AI summaries: Model generates contextual summaries (no more “Agent completed after N steps” fallbacks)
  • 3-tier summary priority: 1) complete_task tool summary, 2) AI’s last text response, 3) descriptive fallback from tool history
  • Custom stop conditions via abort: Consecutive failures, same tool repeated, timeout checked in onStepFinish

Improvements Applied (v2)

Performance

  • Parallel execution: ./run-all.sh --parallel runs all tests concurrently (~3x faster)
  • Timeout handling: All curl calls have --max-time limits (30s API, 90s agentic, 120s per-test)
  • Per-test timing: Results table shows duration for each test
  • Suite timing: Total wall-clock time reported at end

Reliability

  • agentic_request() helper: Centralized SSE request function with built-in timeout
  • Test isolation: --isolated flag creates fresh conversation per test (no state pollution)
  • Timeout detection: Tests killed after timeout reported as TIMEOUT (exit code 124)
  • Cron test fixed: Updated prompt to name tools explicitly (cron tools now available via getAllChatTools())
  • Web search test fixed: More explicit prompt enforces tool chaining

Code Quality

  • Reduced duplication: All agentic tests use agentic_request() helper from config.sh
  • Reusable SSE parser: parse_sse_stream() and check_expected_tools() in config.sh
  • Timing helpers: now_ms() and format_duration() (macOS compatible)
  • New CLI flags: --parallel, --isolated, --verbose, --timeout N

Key Findings

Working Well

  • Multi-step tool chains work correctly (project -> ticket -> update -> get)
  • Parallel tool calls in single step (git_status + git_log)
  • Error recovery: model retries with different approach after tool failure
  • AI SDK v6 native multi-step with stopWhen + onStepFinish
  • Cron tools accessible in agentic mode (getAllChatTools fix confirmed)
  • Proper AI-generated summaries with ticket links, step details, and context

Known Behaviors

  • Model sometimes uses web_fetch instead of web_search (gpt4o-mini preference)
  • Cron trigger may return “no job found” if job name doesn’t match exactly

Test Environment

  • Server: pnpm dev running on localhost:3000
  • Model: Azure GPT-4o (fallback from Anthropic - no ANTHROPIC_API_KEY set)
  • PROFCLAW_MODEL=gpt4o-mini (default in config.sh)
  • Conversation persistence: Working via .test-state.json

Usage

Test Scripts