Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills
NVIDIA DOCA Agent Skills help developers build applications on BlueField DPUs faster. This article explains what the agent skills provide, how they fit an AI agent workflow for DPU development, and where practical examples can shorten the path from setup to a working BlueField application.
Tags
Quick summary
NVIDIA DOCA Agent Skills help developers build applications on BlueField DPUs faster. This article explains what the agent skills provide, how they fit an AI agent workflow for DPU development, and where practical examples can shorten the path from setup to a working BlueField application.
Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills
Building on a DPU is not the same as building on a host CPU, and the difference shows up early. A DOCA application usually splits across two execution domains: a host-side control plane that configures and supervises, and DPU-side logic that runs on the BlueField's Arm cores and touches the networking datapath. Getting an AI coding agent to help with that split is harder than pointing it at a single-process codebase, because the agent has to hold device context, SDK conventions, and a build-and-deploy loop in its head at once.
NVIDIA's developer blog post Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills (published October 1, 2026) addresses that friction directly. This article works from that single primary source. Where I describe workflow mechanics that the post itself does not spell out, I label them as interpretation or as standard DOCA practice rather than as claims from NVIDIA.
What the Source Establishes
The verified facts are narrow, and it is worth stating them plainly before building anything on top:
- NVIDIA published a developer blog post titled "Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills."
- The post is available at the NVIDIA Developer blog.
- The subject is accelerating application development on NVIDIA BlueField using NVIDIA DOCA Agent Skills.
That is the factual floor. The post is the canonical reference for the specific skills NVIDIA ships, the exact formats involved, and the release versions those skills target. Everything below that describes how to work with such a setup is my synthesis of general DOCA development practice, clearly marked as interpretation where it goes beyond the announcement.
Why BlueField Development Resists Naive Automation
Interpretation follows, not a claim from the source.
A generic coding agent does reasonably well on a self-contained library with a well-known API. BlueField work breaks three of the assumptions that make that possible.
First, two domains, one feature. A flow-offload feature is rarely complete on one side. Something on the host must discover the device, open a DOCA context, and push configuration. Something on the DPU must react to that configuration in the datapath. An agent that generates only the host half produces code that compiles and does nothing useful.
Second, hardware capability is a runtime fact. Whether a given offload, encryption mode, or queue primitive is available depends on the specific BlueField generation, the firmware, and the DOCA release installed. Code that assumes a capability is present will fail at runtime rather than at compile time, which is exactly the failure mode an agent is worst at diagnosing without feedback.
Third, the build loop is not `make && ./run`. You typically cross-compile or build on the DPU, deploy the binary to the Arm side, and then observe behavior on a second system. Each extra hop is a place where an agent's assumptions about the filesystem, the toolchain, or the target architecture can silently diverge from reality.
Agent Skills matter precisely because they attack this context problem rather than the code-generation problem. The value is not a smarter model; it is a narrower, better-grounded one.
What "Agent Skills" Change in Practice
This section is interpretation, framed as a working model rather than a description of NVIDIA's implementation.
The useful mental model is that a skill is a packaged unit of domain procedure: when to use it, what preconditions to check, which commands or APIs to prefer, what the failure signatures look like, and how to verify success. It is closer to a runbook an experienced engineer would hand a new colleague than to a prompt snippet.
Applied to BlueField, that means a skill can encode the parts of the workflow that a model cannot infer from the code alone:
- The discovery sequence — how to confirm a DPU is present, which tools report capability, and what "expected" output looks like.
- The domain split — which parts of a feature belong on the host and which belong on the DPU.
- The verification contract — what a passing test looks like for this class of offload, and what a silent misconfiguration looks like instead.
- The environment invariants — the OS, kernel, and DOCA version assumptions under which the procedure is known to work.
The practical consequence is a shift in how you spend review time. Instead of reading every line of generated code hunting for hallucinated API calls, you spend it checking whether the agent respected the skill's preconditions. That is a much smaller surface.
Requirements
The requirements below are generic to a DOCA development environment. Exact versions come from NVIDIA's DOCA documentation and the blog post, which are authoritative for your release.
- A host system with an NVIDIA BlueField DPU installed and visible on the PCIe bus.
- A supported Linux distribution on the host, and a kernel matching what your DOCA release requires.
- Access to the DOCA software stack, installed either from NVIDIA's package repository or from the relevant DOCA installer, following NVIDIA's official installation instructions.
- A toolchain: a C/C++ compiler,
cmake,pkg-config, andgit. - Python 3 with
venvif you intend to script the host-side loop or run an agent harness locally. - Network access for package installation, and appropriate permissions — most of the discovery and installation steps below need
sudo.
One caution worth stating explicitly: do not hardcode package names or repository URLs from a blog post into a provisioning script. Package naming changes between DOCA releases. Discover what your environment actually offers instead.
Step-by-Step Installation
The sequence below is deliberately discovery-first. Every step asks the system what it has before asserting anything about it.
1. Confirm the BlueField device is present
List NVIDIA/Mellanox PCIe devices on the host.
lspci -nn | grep -i mellanoxYou are looking for a network controller entry. A BlueField DPU presents itself as a PCIe device with its own Arm subsystem, so the exact function layout will differ from a plain NIC. If nothing appears, stop here — no amount of SDK installation will help, and the agent will be working against a phantom device.
2. Record the host environment
Capture the OS and kernel so you can pin your skill definitions to a known-good combination.
cat /etc/os-release
uname -srmKeep this output. When the agent produces code that fails only on your machine, this is the first thing you compare.
3. Check whether DOCA is already installed
Query the package database rather than assuming a clean system.
dpkg -l | grep -i doca || echo "no doca packages found via dpkg"Then look for an installation directory, without guessing the exact path used by your release.
find / -maxdepth 4 -type d -name 'doca*' 2>/dev/null4. Discover the packages your repository exposes
This is the step that replaces guessed package names.
apt-cache search --names-only '^doca' | sortRead the list. It tells you which DOCA components your configured repository carries and, indirectly, which release line you are on. Cross-check it against NVIDIA's official DOCA installation guide before proceeding.
5. Install the DOCA components you need
Update the index, then install the components identified in the previous step. The placeholder below is intentional — substitute the real name from your own apt-cache search output.
sudo apt-get update
sudo apt-get install -y <doca-package-from-step-4>6. Install the build toolchain
The compiler and CMake are needed for any DOCA application you or the agent will write.
sudo apt-get install -y build-essential cmake git pkg-config7. Locate the DOCA query and sample binaries
Instead of assuming a tool name, ask the installed package what it put on disk.
dpkg -L <doca-package-from-step-4> | grep -E '/(bin|sbin)/'This is how you find the version-query utility and the sample applications for your release. Sample applications are the single best grounding material you can hand an agent, because they are guaranteed to compile against your installed headers.
8. Create a project scaffold
Set up a predictable layout so the agent has a stable place to write and read.
mkdir -p ~/bf-doca-lab/{src,skills,scripts,build}
cd ~/bf-doca-lab
git init9. Prepare a Python environment for host-side scripting
If you plan to script discovery, deployment, or agent invocation, isolate the dependencies.
cd ~/bf-doca-lab
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pipUsage Examples
The examples below are patterns, not a specification of any NVIDIA-provided format. Treat them as scaffolding to adapt.
Example 1: Encode a skill as a checked procedure
A skill in this general shape gives the agent preconditions and a verification contract, not just instructions. Save it under skills/.
# skill: dpu-device-preflight
## Preconditions
- Host has a BlueField DPU visible on PCIe
- DOCA packages installed and query tools present on PATH
## Steps
1. Confirm the DPU is enumerated on the PCIe bus.
2. Record the DOCA release for this environment.
3. Confirm the required capability is reported by the query tool.
4. If the capability is absent, STOP and report it. Do not generate code.
## Verification
- Every step produces observed output, not an assumption.
- Absence of a capability is a valid, reportable result.The "STOP and report" clause matters more than it looks. It converts a class of hallucination — inventing an offload that the hardware does not support — into a clean failure.
Example 2: A deterministic build-and-test loop
Give the agent a single command that either passes or fails, so it cannot narrate its way around a broken build.
#!/usr/bin/env bash
set -euo pipefail
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j"$(nproc)"
ctest --test-dir build --output-on-failureBecause set -euo pipefail is in effect, any failing stage aborts the script. The agent sees a non-zero exit code and the tail of the log, not an ambiguous success.
Example 3: A host-side orchestrator that logs observed state
This template runs an external command, captures its output, and logs both the command and the result. Adapt the command list to the tools you found in step 7 of the installation.
import logging
import subprocess
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s",
)
def observe(cmd: list[str]) -> str:
"""Run a read-only discovery command and return its stdout."""
logging.info("running: %s", " ".join(cmd))
result = subprocess.run(
cmd,
capture_output=True,
text=True,
check=False,
)
if result.returncode != 0:
logging.warning("command exited %d: %s", result.returncode, result.stderr.strip())
return result.stdout
if __name__ == "__main__":
for probe in (
["lspci", "-nn"],
["uname", "-srm"],
):
output = observe(probe)
print(output)The read-only constraint is deliberate. Discovery commands are safe to run unattended; configuration commands are not, and a skill should separate the two phases explicitly.
Example 4: A prompt pattern that keeps the agent grounded
When you invoke the agent, provide the skill file, the environment snapshot from step 2, and the raw failure output. A compact pattern:
Context:
- Skill file: skills/dpu-device-preflight.md
- Environment: <paste /etc/os-release and uname output>
- Installed DOCA release: <paste version output>
Task:
Diagnose why the build in scripts/build-and-test.sh failed.
Propose a patch only for the host-side component.
Do not introduce API calls that are not present in the installed headers.
Required output:
1. Root cause, with the log line that supports it.
2. A unified diff.
3. The exact command to verify the fix.The final line is the important one. Requiring a verification command closes the loop and makes the agent's claim testable.
Open Limits and What to Verify Yourself
Three things remain genuinely uncertain, and pretending otherwise would be misleading.
The specifics of NVIDIA's skills. The blog post is the authoritative source for which skills exist, what format they use, and which releases they target. This article does not reproduce that inventory.
Version drift. DOCA package names and tool locations change between releases. The discovery-first installation sequence above is designed to survive that drift, but it cannot substitute for reading the release notes.
Performance claims. The title says "faster," and the source frames the work as accelerating development. I have not asserted a speedup figure, a percentage, or a benchmark, because none was provided to me. If you need a number for a decision, run your own measurement: pick a representative DOCA feature, build it once with your existing process and once with the agent-plus-skills workflow, and compare engineering hours to a working, verified implementation. Elapsed model time is not the metric that matters here.
Conclusion
The interesting thing about DOCA Agent Skills is not that an agent writes DOCA code. It is that the bottleneck in DPU development has always been context — device capability, the host/DPU split, and a build loop that spans two systems. Skills package that context so an agent can be useful within it rather than plausible outside it.
The practical path is short. Confirm the DPU is on the bus, discover what your DOCA installation actually provides instead of assuming, give the agent the sample applications and a deterministic build-and-test script, and require a verification command with every proposed change. Then measure the result against your own baseline. That is how you find out whether "faster" is true in your environment, which is the only environment that counts.



