An AI feature is not tested until the real response streams
A clean build and mocked tests prove the application fits together, but not that the live model, credentials, and streaming path work.
The interface compiled. The automated tests passed. The agent still could not answer because the running environment did not have the required credential. After the credential was added, the process also needed a restart before the real streamed response could be verified.
That is why I now separate three claims: the code builds, the behavior works with controlled test data, and the full live path works with the actual service. They are all useful. They are not interchangeable, however enthusiastically the test runner uses green.