What do the first 30 days of a Claude implementation look like?
Week 1 — discovery and scoring. We sit with the teams who own the workflow, score candidate use cases on value, data readiness and risk, and agree the measurable outcome the engagement is judged on. You leave the week with a shortlist and a target metric, not a deck.
Week 2 — eval suite and access. We assemble a golden dataset from your real cases, write the graders, and get model access wired through Bedrock, Vertex or the Anthropic API inside your network boundary. From this point every change is measured rather than argued.
Weeks 3 and 4 — working prototype. We build the prompt, retrieval and tool layer, stand up an MCP server over the systems the workflow touches, and run it against the eval suite until it clears the agreed bar. You see a system handling your own cases end to end.
End of month one — a go or no-go decision backed by numbers: accuracy against the eval set, cost per interaction, latency, and the remaining work to reach production. If the numbers do not support building, we say so.
