Skip to main content

Atlas Agent Engine: Common Problems and Errors

This article provides initial steps for common Atlas Agent Engine setup, build, deployment, invocation, connectivity, memory, and workflow issues.

P
Written by Pavan

Please do not share passwords, tokens, API keys, secret values, connection strings, private endpoints, or sensitive data while reaching out to MongoDB Atlas chat support.

Quick troubleshooting reference

Issue or error

What it means

How to resolve

Documentation

I receive a 401 Unauthorized or an authentication failure.

The request does not have a valid authenticated session for the intended environment.

Complete the supported sign-in flow, select the intended context, and retry.

I receive a 403 Forbidden, “access denied,” or “not authorized” error.

The selected identity does not have access to the selected resource or action.

Verify the organization, project, workspace, and role, then start a new authenticated session if access changed.

My agent reports a missing secret, configuration value, or a credential error.

A required runtime environment variable is absent, or its configured name does not match.

Check the secret declaration and configure the matching cloud secret; do not expose secret values.

I enabled memory and now my agent fails or behaves unexpectedly.

The memory feature, or one of its required prerequisites, may need configuration.

Verify the memory configuration, required secret, and Atlas setup, then deploy the updated configuration.

My agent stopped working after a configuration, SDK, or dependency change.

The updated manifest, dependency set, or deployed image needs validation.

Validate the manifest, restart locally after dependency changes, or rebuild and redeploy for a deployed agent.

Before you begin

Identify the stage at which the issue occurs: sign-in, local run, build, deployment, invocation, tool or external-service connection, memory, or workflow configuration. Keep the timestamp and non-sensitive error details ready as you follow the guide.

1. My local agent will not start or does not behave as expected

A local startup or behavior issue usually requires checking the local project directory, container runtime, environment file, and running local services. Verify each of these before changing the agent implementation.

Steps to troubleshoot

  1. Start from the root of the agent project, where agent.yaml is located, and follow Run agents locally. Meaning that the locally deploying on Docker and test before deploying on Agent Execution Runtime (AER).

  2. Confirm that the local container runtime required by the documentation is running before starting the local environment.

  3. Use the local status command in Run agents locally to check the local services. Review the relevant local service output with the documented `local log` command.

  4. Check the local .env file used by the project. Correct the file if it contains an invalid environment entry, and keep credentials in the environment file rather than in the source code.

  5. If you changed the Python dependencies, synchronize the dependencies and use the documented restart workflow in Run agents locally. If you changed the TypeScript dependencies, follow the documented local rebuild or restart workflow.

  6. Run a test, known request using Test your agent after the local stack is healthy. This separates a local startup issue from the behavior of a specific agent request.

2. My agent build failed or I received an image-build error

An agent build or image-build failure means the project configuration, manifest, dependencies, or selected runtime must be checked before creating another build. Validate the manifest and build inputs first.

Steps to troubleshoot

  1. Confirm that agent.yaml is in the expected project location and that the project layout matches the selected workflow: single-agent or monorepo.

  2. Validate the manifest locally before building with agentengine agent validate. The validation command checks the manifest against the platform schema and catches invalid values before the build pipeline runs.

  3. Check the required entrypoint value in the Agent Contract Reference. It must identify the application object in the documented module.path:attribute format.

  4. Confirm that the declared runtime language matches the project. TypeScript agents must set the documented TypeScript language value; otherwise the build uses the Python runner-base and the agent does not start.

  5. If the project uses private package registries, confirm that the registry declaration, project tooling, HTTPS registry URL, and build secrets use Artifact Registry Validation.

  6. Correct the reported configuration or dependency issue, then create one new build and review its result before making additional changes.

3. My agent build completed, but the deployment failed

A completed build confirms that the container image was built; it does not confirm that the agent deployment and its runtime requirements are satisfied. Verify the deployment status, configuration, secrets, and required services.

Steps to troubleshoot

  1. Confirm the workspace that received the build and review the deployment status using Deploy Your Build.

  2. Keep build and deployment results separate. A successful build confirms that the image was built; any deployment still requires the runtime configuration and required services to become ready.

  3. Review the failed deployment details and logs for the first reported condition or message.

  4. Confirm that the secrets mentioned in the agent’s configuration exist in the target scope. Configure secrets through Provision Cloud Secrets; do not add secret values to the manifest or logs.

  5. If the agent requires Atlas or another external service, complete Set up Atlas resources and Manage network egress policies before retrying.

  6. After correcting the documented prerequisites, deploy the selected build again and wait for the deployment result.

4. My agent works locally but fails after deployment on Atlas Agentic Engine (AAE)

When an agent works locally but fails after deployment, compare the local and deployed configuration. The deployed workspace requires its own runtime configuration, cloud secrets, Atlas connectivity, and network access.

Steps to troubleshoot

  1. Check if the target workspace is available within the target project.

  2. It can be an atlas project where the user is not the project owner. Verify the user role in Atlas and check billing or egress IP’s.

  3. Compare the deployed agent configuration with the current agent.yaml and the Agent Contract Reference. Check the entrypoint, runtime language, enabled features, and required secret names.

  4. Configure the required cloud secrets for the deployed environment with Provision Cloud Secrets. Local .env values do not automatically configure the deployed workspace.

  5. Confirm that the deployed agent can reach its required Atlas cluster and external services. Apply Manage network egress policies to the deployed environment, not only to the local machine.

  6. Send a test request using Invoke an Agent once deployment and configuration are complete.

  7. If the local and deployed paths use different agent code, dependencies, or configuration, rebuild and redeploy the corrected project before testing again.

5. My agent does not respond, invocation fails, or a request times out

An invocation failure, no response, or timeout requires confirming that the workspace is healthy and that the request uses the documented invocation workflow. Then verify dependencies such as provider credentials, tools, and network access.

Steps to troubleshoot

  1. Confirm that the workspace deployment completed successfully before troubleshooting the request itself.

  2. Use the Invoke an Agent for the documented invocation method and request format for the deployed agent.

  3. Start with a test request that does not depend on a large input, external tool, or long-running workflow. Confirm whether the same request succeeds.

  4. Check that the configured LLM provider credentials are present in the deployed environment and that required tools or external services are available.

  5. If the client needs incremental output, use the streaming invocation workflow in Invoke an Agent and ensure that the client processes stream completion and error events.

  6. Review Monitor agents and deployments and Deploy Your Build before retrying after a configuration change.

6. My agent cannot reach an external service, tool, or remote MCP server

An external-service, tool, or remote MCP connection failure requires checking the configured destination, the deployed network-egress policy, and the credential configuration for that integration.

Steps to troubleshoot

  1. Identify the external host and port required by the agent, tool, or MCP server without copying tokens, authorization headers, or private request data.

  2. Add the destination through Manage network egress policies for the agent’s deployed environment. The network policy controls outbound access for the configured sandbox.

  3. Confirm that the external service also accepts traffic from the deployed environment and Configure any required authentication credentials through Provision Cloud Secrets.

  4. For a remote MCP server, confirm that its agent.yaml configuration follows Use Remote MCP Servers.

  5. For MCP OAuth, complete the local authorization flow in Use Remote MCP Servers, check the cached credentials, and upload the credentials for the deployed workspace before retrying.

  6. Invoke a test request that calls the relevant tool after the destination, credentials, and egress policy are configured.

7. My deep-agent or multi-agent workflow is not working

A deep-agent or multi-agent workflow issue requires validating the feature configuration, supported workflow shape, and the manifest for every participating agent before testing the full workflow.

Steps to troubleshoot

  1. Confirm that the project follows Build a Deep Agent or Agent-to-agent communication, as applicable.

  2. For a deep agent, enable the deep-agent feature described in Build a Deep Agent in agent.yaml before building the agent.

  3. Confirm that the agent uses the documented platform API for the deep-agent workflow and that the entrypoint and supported framework configuration match the Agent Contract Reference.

  4. For a multi-agent repository, use the monorepo layout in the agent contract reference: a root agent.yaml with an agents list, plus a single-agent agent.yaml in each agent subdirectory.

  5. Validate each agent manifest, then build and test the relevant agent or workflow with a test request before testing the complete workflow.

  6. Ensure that each participating agent has the documented secrets, network access, and deployment configuration required by its own tools and integrations.

Official documentation

Did this answer your question?