Elevate Your Content Creation with the Codility MCP Integration
What this lets you do?
Describe a task in your AI coding assistant, and the agent builds the full Codility task for you: the task itself, the test cases, the golden solution, and the candidate stub. The agent validates the output end to end, then publishes the task to your Codility Task Library. Most tasks take about ten minutes from describing to published.
You stay in the IDE your team already uses. No web editor, no back and forth, no copy-paste. It works for both candidate environments, the Basic IDE and the full VS Code IDE. You pick the environment up front, and the rest of the flow is the same.
What you get?
- A complete, validated task. Multi-file structures supported. Test cases included. Golden solution scores 100%, candidate stub scores 0%, or the agent self-corrects until it does.
- Screen or Interview task types. Same flow for both.
- Your content, not a template. Point the agent at your real codebase, a pull request, or a job description, and the task reflects your stack.
Who it's for?
Anyone on your team with access to Codility who wants to create custom assessment content. You do not need to write code to use it, but you'll get the best results if the person describing the task has engineering context on what they want to assess.
Before starting you'll need:
- A Codility account with admin access. Build with MCP currently requires admin permissions to publish to your Task Library.
-
An MCP-compatible AI coding assistant. Any of these works:
- Claude in VS Code (richest integration, recommended starting point)
- Cursor
- GitHub Copilot in VS Code
- Any other MCP-compatible client
- About three minutes for setup, the first time.
If you're not sure whether your AI assistant supports MCP, check the assistant's documentation for "Model Context Protocol" or "MCP server."
Setup
1. Open the Task Library
From the Codility main menu, go to Tasks, then Create tasks. Click the blue + Create task button, then click Build with AI.
2. Select the environment for the task
Pick where the task will run for the candidate. Each environment connects with its own key, so the configuration you copy in the next step belongs to the environment you pick here.
- Basic IDE. Tasks run in the Basic IDE and are scored automatically. Use it for self-contained coding problems, language-specific or technology-agnostic. It is the quickest route to a scored Screen task.
- VS Code IDE. Tasks run in the full VS Code workspace, again language-specific or technology-agnostic. Use it when the task needs a project rather than a single problem: multiple files, a terminal, dependencies, backing services, or a codebase the candidate has to read before changing anything.
If you're unsure: Basic IDE for algorithmic or single-problem tasks, VS Code IDE for anything that should feel like working in a repository.
3. Choose an installation method
Two options:
Option A: One-click install into VS Code Click Install in VS Code. Your browser hands off to VS Code, the MCP server configuration writes itself to your local config, and you're authenticated through your Codility login.
Option B: Choose your AI tool and add the server by copying the configuration Pick your client from the tabs (VS Code, Cursor, Claude Desktop, Claude Code, Codex), click Copy, and paste the JSON block into that client's MCP configuration. For VS Code this is .vscode/mcp.json in your project; each of the other tabs shows the path or command for its own client. The configuration includes a time-limited auth token, so add it while that token is still valid.
3. Verify the connection
Open your AI assistant. Ask it something like: "What MCP servers are connected?" You should see Codility in the list. If you don't, see Troubleshooting below.
That's the setup. You only do it once per environment.
Your first task
The fastest way to get a feel for how this works is to ask for something specific. Open a chat with your AI assistant and try a prompt like one of these:
For a focused Screen task:
Create a Codility Screen task in Python that tests a candidate's ability to debounce async API calls. The task should involve fixing a flawed implementation so that repeated calls within 500ms are deduplicated. Include at least five test cases, one of which covers the edge case of very rapid consecutive calls.
For a multi-file Interview task:
Create a Codility Interview task in Node.js. Give the candidate a small Express app with three routes and one known bug in the authentication middleware. The candidate should identify and fix the bug, then add rate limiting to the login route. No automated scoring, this is for live interview use.
For migrating existing content:
We have an existing custom assessment in CodeSignal that tests Python data pipeline skills. It asks candidates to implement a CSV-to-JSON transformer with error handling. Rebuild this as a Codility Screen task with equivalent test coverage.
The agent will walk you through the build, ask clarifying questions if your description is ambiguous, and run validation checks. You'll see each step as it happens.
When the task is ready, the agent will ask you to confirm before publishing. You can always edit the task in Codility afterward.
Refining and editing
After a task is generated, you have two ways to improve it:
- Ask the agent for changes. "Make the edge cases harder." "Add a test for empty input." "Reframe the task around a user-facing product scenario." The agent regenerates, revalidates, and shows you the diff.
- Edit in Codility directly. The published task is a normal Codility task. Edit the description, test cases, or solutions the same way you would any other task.
Most customers start with agent-driven refinement, then polish in Codility once the structure is right.
Tagging your task
Tags are how a task becomes findable later. They drive the filters in the Task Library, and they are what your colleagues check before deciding whether to write a new task or reuse an existing one. Nobody finds an untagged task.
The MCP server handles tagging for you. It exposes a list-tags endpoint (list_tags on the VS Code server) that returns the current tag vocabulary grouped by axis, so the agent can read the allowed names before it publishes. You do not need to look anything up. Say what the task is about and ask the agent to tag it.
What can be tagged. Tags come back in seven groups:
- Category. What kind of exercise it is: coding, bugfixing, algo, api, real-life, multi-stage, live-interview, ai-native, work_with_ai.
- Technology. The stack the task uses: react, django, spring-boot, postgresql, kubernetes, pytest, terraform.
- Skill. What the task measures, prefixed with skill-: skill-debugging, skill-system_design, skill-code_review, skill-data_querying, skill-ai_output_validation.
- Job Role. Who the task is aimed at: backend-developer, frontend-developer, data-engineer, qa-engineer, devops-engineer, security-engineer, role agnostic.
- Business Context. The domain the scenario is set in: banking-and-finance, e-commerce, healthcare, logistics, gaming.
- Time Limit. The intended length, from 5-minutes to more-than-1-hour.
- Spoken Language. The language the statement is written in: en, de, fr, es, pl, jp.
Each group holds more names than the examples above; list-tags returns the full set.
How to ask for it. Anything along these lines works:
Before publishing, call list-tags and tag this task appropriately: category, technology, the skills it measures, the job role it targets, and a 30-minute time limit.
Add the e-commerce business context and a skill tag for debugging, then re-publish.
Two constraints. The vocabulary is curated, not free text: the agent can only reuse names that already exist, and invented or near-miss spellings are rejected. The programming language is set by the platform from the task itself, so nobody assigns it as a tag. A task can carry up to 20 tags.
You can add or correct tags afterwards in the task editor, or ask the agent to update them.
Keeping your Task Library tidy
Task creation now takes minutes, so libraries grow quickly. A little discipline at publish time keeps yours usable a year from now:
- Tag every task at publish time, not later. Ask the agent to tag as part of the publish step. Tasks left to be tagged later tend to stay untagged.
- Search before you build. Ask the agent to list existing tasks matching the role and technology you have in mind. A variant of an existing task is often the better answer.
- Be consistent across the team. Agree on which job roles and time limits you use, and tag the same way every time. Filters are only as good as the tags behind them.
- Name tasks so a human can scan them. A title that says what the candidate does beats a clever one.
- Retire what you no longer use. Leaked or superseded tasks are noise for everyone browsing the library.
Data and security
What leaves your environment: Only what you put in the agent's context. If you describe a task abstractly, nothing of yours goes anywhere. If you point the agent at your own code, that code is part of the LLM request.
Which LLM runs the generation: Whichever model your AI assistant is configured to use. Codility's MCP server is vendor-neutral. If your assistant uses Claude, the data handling follows your agreement with Anthropic. If you run a self-hosted model, your data stays in your environment.
What Codility stores: The final published task and its metadata, the same way any other Codility task is stored. We don't log intermediate generation requests or the code context you provide to your assistant.
Compliance: Build with MCP is covered by Codility's existing SOC 2 Type II, ISO 27001, and GDPR posture. If your security team needs deeper architecture detail, your CSM can arrange a briefing.
Troubleshooting
The agent says it can't see the Codility MCP server. Restart your AI assistant. Some clients load the MCP configuration at launch only. If that doesn't work, verify the configuration file is in the expected location for your client and that the auth token hasn't expired.
I see "unauthorized" or "token expired." Your auth token is time-limited. Go back to the Build with AI panel in Codility, pick the same environment, and generate a new configuration. The one-click VS Code install handles this automatically.
The generated task fails validation over and over. This sometimes happens on very niche stacks. Try one of these:
- Simplify the description to isolate the core skill you're testing.
- Provide more context. Point the agent at a real code sample or a pull request.
- Ask the agent to produce a smaller version first, then extend it.
The agent rejects a tag I asked for. The tag vocabulary is curated. Ask the agent to call list-tags and pick the closest existing name. If nothing fits, contact your Codility account team.
The agent finished, but the task isn't in my Task Library. Check that you clicked Publish in the final confirmation step. Generated tasks stay in a draft state until you publish.
Setup works in VS Code but not in Cursor. Cursor's MCP support evolves quickly. Update Cursor to the latest version, then re-run the configuration. If that doesn't work, reach out.
Still stuck: Email support@codility.com with the model and client you're using, a description of what you tried, and any error messages.
FAQ
Can I use this without Claude? Yes. Any MCP-compatible AI assistant works. Claude in VS Code has the richest integration, but the MCP protocol is vendor-neutral.
Is there a cost? No additional charge at launch. Build with MCP is included in all existing Codility plans. You do need access to an AI coding assistant, which may have its own cost depending on which one you use.
Does it work for Interview tasks, or just Screen? Both. Same workflow. Choose the task type when you describe what you want.
Can multiple engineers on my team use it? Yes. Anyone with admin access to your Codility account can set up their own MCP connection. Tasks published by one engineer are visible to everyone with Task Library access.
What about non-technical admins who can't describe tasks themselves? Two options. Pair with an engineering partner who describes while you review and publish. Or point the agent at a job description you already have. It can generate a task from a JD without requiring you to know the implementation details.
How does this compare to paying Codility to build a custom task for me? Custom content services still exist for customers who want vendor-built tasks. Build with MCP is the self-serve alternative. Most customers find MCP faster and cheaper, but both paths are available.
What if the LLM gets something subtly wrong? The validation gate catches failures in the logic. It can't catch bad judgment. Review the final task the same way you would review any custom task before sending it to candidates.
More resources
-
3-minute walkthrough video:
Limitations
Task replay: Multi-file freeform tasks generated via MCP for the Basic IDE do not include the code-replay timeline (play, pause, rewind) in the candidate report. The timeline is available for single-file Monaco tasks and for VS Code tasks created through MCP, where it is in Preview for VS Code Native tasks in Interview. Multi-file freeform Basic IDE tasks are not in scope.
Library metadata edits: The MCP flow manages task content (problem statement, code, tests, environment) and tags, which the agent can set at publish time using the list-tags vocabulary. It does not set the remaining library metadata on the task record. Set these in the task editor in the Codility UI after the task is created:
- Difficulty
- Assessment type
Environment and sidecar catalog: The list of supported environments and backing services is presented from a fixed catalog inside the MCP flow. Supported services include PostgreSQL, MongoDB, Redis, MySQL, and Localstack. If you need an environment or backing service not in the list, contact your Codility account team.
Custom Docker images and dependencies: MCP task creation runs against a Codility-maintained set of pre-built environments. Pinning a custom Docker base image, custom system packages, or customer-defined dependency layers is on the roadmap and not available today.
Task types: MCP / Cody creates freeform-style tasks (with optional multi-file). It does not create classic single-file Monaco tasks, hybrid tasks, or pre-built library tasks from the Codility catalog. Use the standard library or the task editor for those.
Bulk operations: MCP / Cody creates one task per session. Bulk import, batch duplication, or programmatic re-publish are not part of the flow.
Feel free to reach out to your Customer Success Manager or support@codility.com if you have any additional questions.
Happy building.