Skip to content

I Gave Hyperagent One Handle and Nothing Else: It Verified the Rest

Not my name, not a CV, not a list of my sites. Six words containing a GitHub handle, typed into an empty account. It came back with seven live subdomains, more than thirty repositories, two certification badges, a Docker Hub namespace and a tutorial on somebody else's platform, every claim graded against an independent source. Twice, a day apart, for $21.02.

The reports are the primary source, not this page

  • first audit · 2 Aug 2026 · read the source code directly and carries the claim-by-claim grading
  • second audit · 3 Aug 2026 · ran against the corrected version

Both threads, showing the single line of input and every tool call that followed, are linked at the end. What follows is what they found, and where they were wrong.

The setup

Hyperagent is an agent platform that runs on its own cloud machine and browser. I came across it through a YouTube walkthrough offering $500 of trial credit, made an account, and typed one line:

do a detailed research on ibtisam-iq

That is the entire input. No name, no links, no CV, no hint about what ibtisam-iq is or whether the person asking has any connection to it. The account was new, with no integrations connected and no profile filled in.

The reason for keeping it to a handle is that everything published about my work was written by me and nobody had ever checked it. Anything added to that prompt would have started shaping the check.

What it had to work out first

It began by identifying whose handle it was. The first report opens by ruling out three other engineers with similar names before settling on the right one. Identification work is not something a system does when it has been told who its subject is.

Then it had to locate the material. Nothing pointed it at nectar, runbook, cert-vault, the projects site, the blog, the DebugBox docs or the portfolio itself. It reached all seven, plus the Credly badges, the iximiuz Labs author page, a KodeKloud verification URL, and a Docker Hub namespace published as mibtisam rather than ibtisam-iq, so not even a matching-string guess.

The estate was discovered, not supplied.

How it actually ran

This is the part worth knowing before trying it. It is not one model answering a question.

Both runs used a fan-out pattern. An Opus 5 orchestrator wrote a plan, dispatched four parallel research clusters as Sonnet 4.6 sub-agents, then collected and synthesised the results. Search went through Exa. Sub-agents carried names and briefs; the two visible in the transcripts are Beacon · Portfolio Analyst and Sprocket · Repo Cartographer.

The clusters split the work: GitHub repositories, the seven self-hosted sites, credentials and professional profile, and external community footprint.

In the first run the orchestrator did not only delegate. It pulled the GitHub and Docker Hub APIs directly, read Dockerfiles, CI workflows, the Jenkinsfile and the Makefile as raw files, and captured screenshots of the Credly badge pages. Its own summary put it plainly: "Four agents swept his sites, repos, docs and external footprint while I pulled the GitHub and Docker Hub APIs directly, so the numbers below are measured rather than quoted."

Run one Run two
Date 2 Aug 2026 3 Aug 2026
Sonnet 4.6 (sub-agents) $8.21 $4.40
Opus 5 (orchestrator) $5.62 $2.58
Exa (search) $0.08 $0.13
Total $13.91 $7.11

Both came out of the trial credit. $21.02 is the entire barrier to entry for two independent audits, which is the actual argument for doing this.

What it verified

Claim Result
CKA and CKAD Live Credly badges, certificate IDs, 2027 expiries
Commit history 2,581 public commits, account created July 2024
DebugBox adoption 1,300 Docker Hub pulls
iximiuz Labs 1 reviewed tutorial, 5 working playgrounds
Project count 8 projects, 75 technologies

That pull count is the only figure in either report produced by other people. Stars and followers can come from goodwill. A pull is somebody with a broken cluster deciding the image was worth trying.

Where my own numbers were wrong

I rounded up. I advertise DebugBox lite as 93% smaller than netshoot. Both images were pulled from the registry and measured: 14.6 MB against 203 MB, which is 92.8%. Two tenths of a point in my own favour.

I rounded down. Four documentation sites advertise "200+ pages". The sitemaps return 665.

Site Pages
nectar 365
runbook 182
cert-vault 81
blog 37
Total 665

I let three pages go stale. A later DebugBox release removed tools from the images, which made all three variants smaller. That went into the project CHANGELOG. It did not go into the portfolio, the projects site or the GitHub profile README, all written before that release.

Variant The pages said The product was at
lite 14 MB 15 MB
balanced 50 MB 47 MB
power 104 MB 91 MB

Small on two rows. Power was out by 13 MB, which is most of a lite image. Fixed at the source rather than page by page, so the figures now come from one place.

A fact with more than one home drifts at every copy

The release was correct and the changelog was correct. The three pages describing the images were a version behind. Nothing failed, nothing logged, and no test covers prose. Any number that appears in more than one place needs a single source or it will eventually disagree with itself, and the copy that goes stale is always the one nobody reads with fresh eyes.

What it got wrong

Three findings arrived with full confidence and were false.

Routes redirecting to Credly. The first run reported that several portfolio routes redirected to a badge page, leaving the site with no working About page. All of them return 200 and render their own content. Nothing redirects anywhere.

A tool count set against a technology count. It read a 66-tool skills list on the portfolio, a 75-technology total on the projects site, and called the gap an inconsistency. Different measures, different labels, counting different things.

A stale "6 Projects" counter. The second run reported it sitting above eight rendered projects. That counter is computed from the same data the page renders, and counting that data gives 8 projects and 75 technologies. It read a filtered view and reported the filter as a defect.

An audit is evidence, not a verdict, and that cuts toward the auditor too.

What I did not change

Both runs landed on the same weak points, and both were right.

Signal Value
GitHub stars, all repos 21
Forks 4
Followers 2
External contributors 0 (Dependabot is the only other committer)
Third party mentions 0 across awesome lists, newsletters, HN, Reddit, Dev.to
DevOps employers on record 0
LinkedIn followers against GitHub 5,311 against 2

No on call rotation, no incident under real stakes, no system inherited from another team. The second report graded all of it as strong-junior to early-mid, with documentation habits ahead of that band and production judgement entirely unproven.

None of that is a defect in a portfolio, so none of it was rewritten. Those are facts about where a career sits 25 months after it started.

Both reports reached the same structural conclusion: the work is further along than the audience is. The second called it under discovered rather than over claimed. That describes a choice I made without quite deciding to make it, since almost all of those 25 months went into building and almost none into telling anyone.

The second run was not equivalent

Same prompt, same account, one day later, and a different piece of work.

The cost table above is the evidence. Run two spent $7.11 against $13.91 on a target that had not got smaller. The first run went to primary sources and cross checked almost everything itself. The second leaned on what was already published, including the corrections from the first pass, without repeating that verification.

The outputs read similarly. The token spend says they were not similar work.

Two agreeing reports are not two pieces of evidence

A second run against the same target will happily restate the first one's conclusions, especially once those conclusions are published where it can read them. It looks like corroboration and costs half as much to produce, because it is doing half the work. Any finding that matters has to be traced back to the run that measured it, and the billing page turns out to be a decent proxy for which run that was.

The source material

Everything above is a summary written by the person being assessed, which is the problem this started with. The reports carry what a summary cannot: claim by claim grading tables, the breakdown of all eight projects, a scorecard across nine dimensions, and each run's list of what it could not access.

All four artifacts are public

The threads show the single line of input and every tool call that followed, so each report can be checked against the process that produced it.

The disagreements between them are more useful than the agreements.