Shamalo

SYS/NAV

SYS/AI

AI isn't an offer. It's the axis.

AI engineering, explained for those discovering the craft

An AI engineer doesn't build AI models. They put agents into a framed work process — one step at a time, a written success criterion, a human review — and they measure what that changes in your numbers.

AI workshop

  • Cursor
  • Claude
  • Anthropic
  • Kilo Code
  • GitHub Copilot
  • Cline
  1. 01

    Context

    The numbered goal and the data, written first.

  2. 02

    Plan

    One phase, verifiable tasks.

  3. 03

    Agent execution

    Executes the plan, not a hunch.

  4. 04

    Human review

    Nothing ships without review.

  5. 05

    Go live

    Live, then we measure again.

SYS/DEF

What exactly is an AI engineer?

Someone who plugs AI agents into a team's real work — measure, produce, decide — without dropping quality or the numbers.

An “agent”, here, is an AI assistant given a precise task, with written rules and a result to hit. Three things set it apart from a chatbot:

  • A numbered goal

    Set before the first line of code.

  • Framed agents

    One step at a time, not a free-form chat.

  • A human review

    Nothing goes live without sign-off.

SYS/CONTRASTE

What an AI engineer is not

The craft sits halfway between the developer, who builds, and the AI researcher, who invents models. Here, we're on the build side.

SYS/OUTCOME

What AI engineering changes for you

Three measured effects on real engagements. The method detail comes right after.

  • 3 wks.

    Speed

    The Workfluence site, SEO-ready information architecture included, went live in three weeks.

    Institut Cassiopée: share of tests above +10%
    Same data as a table.
    Test resultShare
    Tests ≥ +10%60%
    Other tests40%

    15 prioritised hypotheses, 8 months of tests.

  • +120%

    Measurement

    At Cassiopée, conversion rose 120% because everything was measured before it was optimised.

    What the loop produced, engagement by engagement
    Measured gains by engagement, as a percentage.
    EngagementMeasured gain
    Cassiopée (UX/CRO)+120%
    Cassiopée (SEA)+65%
    SITL+66%
    Infirmière Reconversion+150%

    SITL: form completion, 27% → 45%. Other rows: conversion rate.

  • 500 pages

    Cost

    Workfluence: 500 job pages produced and reviewed under agents, for about $20 of tools per month.

    • 500

      pages, all reviewed

    • ~$20

      monthly subscription

    • 3

      weeks to live

SYS/WORKFLOW

How an AI engineer works, step by step

We never hand “the project” to an agent. We hand it a step, with its limits and its success criterion. That playbook has a name: GSD. Here is who does what, and in what order.

01 — Protocol

One step at a time, never a vague request

Every step follows the same sequence, in five beats. If verification fails, we redo the step: we never let it pass as “good enough”. That same protocol produced RankyDocky, the studio's internal tool, and the site you are reading.

  1. 01Context — the goal and the limits, in writing.
  2. 02Research — the agent explores what already exists.
  3. 03Plan — tasks, each with a success criterion.
  4. 04Execution — the agent follows the plan, nothing else.
  5. 05Verification — a human reviews before go-live.
ContextResearchPlanExecutionVerificationhumanagenthumanagenthuman
The GSD protocol: five phases in a loop. The token circulates; the active phase lights up. Each cube says who owns the step.
HumanAgentScopeReviewDecideSign offResearchProduceCodeWritebriefdeliverable
The human sends a brief; the agent returns a deliverable. The ping-pong continues until sign-off.

02 — Roles

The human stays in the loop

The agent produces. The human scopes, reviews, decides, signs off. Nothing goes live without that review — that is the whole difference with a chatbot left to “write instead”. Project rules are written once in a file the agent reads every time; it applies them, it doesn't invent them.

  • A written brief, never dictated in a chat.
  • A named owner for every deliverable.
  • Final sign-off is human. Always.

03 — Examples

Two real engagements, from request to deliverable

Same protocol, two very different outcomes: on one side, produce at scale; on the other, call a decision. Under each diagram, the caption says who owns the step.

Workfluence — 500 job pages

Job dataAgent500 pagesLivehumanagentagent × 500 · all reviewedhuman
The client's job data feeds an agent. It produces 500 pages; the bar sweeping the stack is the human review, page by page, before publication. Result: 4,000 visits a month.

Workfluence case →

CRO — from hypothesis to decision

ABShipDropcurrent pagevariantgapmeasuredgapsignificant?gain provengain unproven
Two versions of the same page run in parallel. If the measured gap is not statistically solid, we drop it — that is what happened on the airline A/B test, where the variant pulled down the recap.

Cassiopée case →

SYS/BOUCLE

The Shamalo loop: measure, ship, measure again

The protocol above is the backbone. Here is how it translates on a client engagement, in five beats — and who holds the pen each time.

  1. 01 — Measure

    Set the KPI before the décor

    Human. We look first at the raw data and the conversion funnel — the path a visitor takes through to purchase — then we set the numbered goal. No agent before that. On an A/B test run for an airline, the recap dropped: decision not to ship.

  2. 02 — Instrument

    Make the result readable

    Human, tooled. A tagging plan: the written list of what we want to count (clicks, forms, purchases) and how to count it in GA4, then readable dashboards. At Cassiopée, that foundation is what lets us attribute the +120% conversion to the right cause.

  3. 03 — Agents

    Put the agents in the loop

    Agent, under protocol. Cursor, Claude Code and Copilot execute framed steps — never a free-form conversation. Workfluence: 500 pages produced, then reviewed one by one.

  4. 04 — Ship

    Deliver to production

    The agent produced, the human signs off. The deliverable actually goes live, not into a mockup. RankyDocky: 33 modules, a CRM where each client's data stays isolated. BAO: 34 screens and 10 PRDs.

  5. 05 — Review

    Check, then start again

    Human. Statistical checks, reading by segment, a written recap. For SITL, form completion went from 27% to 45%. Then we run the loop again.

See the loop, step by step →

SYS/STACK

What tools does an AI engineer work with?

One person holds a team's pace because the work is tooled end to end — agents, protocol, written rules — not because a language model “writes instead”.

GSD, the orchestration protocol

The problem with a classic AI assistant: the longer the conversation, the more it forgets and the more quality drops. GSD fixes that by giving each agent a short mission and a fresh memory. Four agents research at once, each on one angle; the plan is reviewed; execution goes out in waves — anything with no dependency moves in parallel; a last agent verifies, and the human signs off.

This is not a tool you will have to learn: it is the studio's internal kitchen. RankyDocky, the internal tool, and this site run on it. The SEO & GEO audit used on engagements comes from it.

ContextStackArchFeaturesRisksReviewed planWave 1Wave 2Wave 3human · discuss4 researchers in parallel
GSD orchestration: human context, fan-out research, a reviewed plan, parallel execution waves.
  1. 01

    Discuss

    The human frames the step: numbered goal, constraints, success criterion. Nothing starts without that.

  2. 02

    Research

    Four agents read the code, the architecture, the features, the traps — in parallel, isolated contexts.

  3. 03

    Plan

    A planner writes verifiable tasks. A checker reviews before the first line of code.

  4. 04

    Execute → verify

    Tasks with no dependency start together. A verifier reviews. The human signs off.

The alternatives to GSD

Other protocols make the same bet: write precisely what you want before letting an AI code. Here are four. The studio chose GSD because the hard part, in practice, is not writing the spec: it is executing several steps in parallel without losing the thread, then verifying everything.

This tab is technical: none of this is on you. You can skip it and miss nothing.

GSDSpec KitOpenSpecTaskmasterDiscussResearch ×4PlanWavesVerifyConstitutionSpecPlanTasksCodeExisting codeDeltaPRPRDGraphTaskTaskTask
Four protocols, four ways of slicing the work. The GSD tile is the studio's: wave execution, isolated context.
  • Spec-first

    GitHub Spec Kit

    The spec drives the code: a project constitution, then the specification, the plan, the tasks, the build. Ideal from a blank page. Less tooled for running several workstreams in parallel.

  • Existing product

    OpenSpec

    You describe only the delta against existing code, not the whole system. Useful for evolving a live product without rewriting all of its documentation.

  • Task graph

    Taskmaster AI

    The product document becomes a tree of tasks, executed one after another. Strong on breakdown, lighter as soon as several agents need to work together.

  • Agile roles

    BMAD

    BMAD: one agent = one role (PM, architect, dev, QA). More ceremony, more “virtual team” coverage. GSD stays thinner: phases, not personas.

Project rules, written once and for all

Every project has its rules: how to isolate a client's data, how to name a page, what an approved screen looks like. They are written once in a file — the studio calls it a “skill” — that every agent reads the same way. No need to repeat the instruction in every exchange: it lives in the project.

You have nothing to maintain. This library is shared across the studio's projects, and every engagement that goes well enriches it, instead of staying in one person's head.

Project rulesRequired frameCompliant deliverableone file, versionedwhat the agent proposeswhat passes the framereviewed · signed off
Project rules are written once in a file. They form a funnel: everything the agent proposes must pass through there before it becomes a deliverable.

Two years with agents in production

Since 2024, AI has been in the daily work, not in the sales pitch. Data analysis, test campaigns, sites and products: the protocol does not change, the tools do. Cursor and Kilo Code hold the workshop; Claude Code leads the execution waves; Copilot completes; GSD frames the whole. The tool list will evolve. The method will not.

202420252026nextdata · dashboardsCRO · SEO · SEAproducts · AI auditssame protocol
Four moments, one protocol. The cases change; the loop discuss → research → plan → execute → verify stays.

SYS/FAMILLE-01

AI engineering applied to growth

The loop measure → form a hypothesis → test → decide runs faster. CRO, organic search, paid ads and measurement become one system instead of four separate workstreams.

Growth & Performance → Cassiopée case →

SYS/FAMILLE-02

AI engineering applied to digital products

From the spec to a clickable prototype, then to live software: deliverables in weeks, not quarters.

Digital product creation → BAO case →

SYS/FAQ

Frequently asked questions on AI engineering

Do the agents work on their own?

No. The GSD protocol frames every step: written context, reviewed plan, execution, human verification. The agent speeds up execution; it never has the last word. The diagrams above show who owns which step.

Without that frame, an agent drifts: it invents a rule, skips a constraint, delivers something that works technically but misses the goal. Review is not a luxury — it is what separates a studio from a text-production machine.

In practice: the human writes the brief and accepts the plan; the agent produces; the human reviews the deliverable before anything goes live. If a step has no named owner, we don't start it.

Is an AI engineer a developer?

In part. It is a profile halfway between developer and analyst: they code, they set up measurement, they can read a conversion funnel and a statistical test — and they steer agents rather than writing everything by hand from start to finish.

It is not someone who merely prompts ChatGPT well, nor a researcher who trains models. The craft, here, is to take a project through to a measurable result: a new page version live, a clickable prototype, a dashboard that answers a question.

If you only need a developer to work through a ticket list, an in-house team or a services firm will do better. If you need someone who connects measurement, product and agents, that is this studio.

Do you train models?

No. This is neither research nor data science in the academic sense: we don't build a homemade AI model. We assemble existing models — from Google, OpenAI, Anthropic — and analysis tools into a process that produces a deliverable.

The value is not “AI”. It is the protocol: a fresh memory at every step, several research tracks in parallel, a reviewed plan, execution in waves, a verification. A model without a frame produces text; a protocol produces a site, a test, a screen.

If what you need is a model trained on your own data — fine-tuning — that is not the offer. We can, however, frame how agents rely on your data, locally, without sending it just anywhere.

Do my data leave my premises?

Your data extracts are analysed locally, on a studio machine, with classic analysis tools. No client data is sent to an AI model without prior written agreement. The perimeter — which files, which columns, which tool — is set at kickoff.

That is deliberately stricter than “we'll paste it into ChatGPT”. A CRM export, a conversion file or a navigation log has no business in a consumer tool by default. If a cloud agent is useful, we say so, we limit the columns, we anonymise what needs it.

You remain the owner of the data and the deliverables. The studio trains nothing on your files and does not reuse them on another client.

Why a one-person studio?

Because instrumentation replaces coordination. One counterpart, tooled agents and a GSD protocol hold a small team's pace, without the cost or the delays of a multi-layer structure.

The hard part is no longer writing the spec: it is executing in parallel and verifying. That is why GSD launches several research agents at once, chains execution waves and requires a verification — rather than adding people and meetings.

There is a ceiling: an international programme with twenty contributors is not the right case. A conversion funnel, a product, an internal tool, yes. The call is there to see which side you are on.

What is GSD, and why not Spec Kit or OpenSpec?

GSD is a lightweight orchestrator: a fresh memory at every step (to keep the conversation from degrading), several research tracks in parallel, a reviewed plan, execution in waves, a verification. It is not one more tool for you to learn — it is the studio's internal kitchen.

GitHub Spec Kit, OpenSpec, Taskmaster or BMAD do something else, and they do it well: spec first, delta against an existing product, task tree, one agent per role. They help write the plan. After two years of real use, the real difficulty is elsewhere: execute in parallel and review — hence GSD.

You don't have to pick the tool. You see deliverables and a cadence. The detail on alternatives is in the “Alternatives” tab higher on this page.

What are “project rules” for?

It is a simple rules file the agent must follow for a given type of deliverable: a conversion audit, a tagging plan, a prototype screen. The studio calls it a “skill”. Without it, the agent improvises; with it, it repeats a gesture already proven.

Studio projects share a library of these files. That is what avoids reinventing “how we write a test” on every engagement. It does not replace review: it makes execution consistent.

You have nothing to maintain. You benefit from the fact they exist — and every engagement that goes well enriches one, instead of staying in one person's head.

Which tools do you use?

For agents: Cursor, Claude Code, Kilo Code, Copilot, Cline, the GSD protocol and project rules. For analysis: Python and SQL. For measurement: GA4, GTM and Looker Studio. For pages: HTML, Tailwind, sometimes Webflow.

The list evolves; the protocol stays. We don't impose our tools: we adapt to what you already have — your CMS, your database, your ad account. The tool serves the goal, not the other way around.

If a tool is a lock (licence, data that cannot leave), we treat it in the scoping. No surprise on day 3.

How long before we see an effect?

It depends on the deliverable, not on “AI”. A clickable prototype: a few days to a few weeks. A first readable conversion test: often a few weeks, the time to accumulate enough visitors. A set of well-ranked pages: months. Usable live software: a few weeks of development after a scoped prototype.

Agents speed up production, not the pace at which your traffic arrives nor the time a test takes to become reliable. Promising a Google #1 in fifteen days because we write faster would be a lie.

The first visible effect, though, is fast: a written scoping within 48h after the call, with the problem, the KPI and the first step.

Do you step into a project already underway?

Yes — that is even the most frequent case. A conversion funnel leaking visitors, a GA4 setup nobody trusts, a mockup never tested with users, a product roadmap stuck. We start from what exists; we don't ask you to start over.

AI engineering then serves to accelerate one piece of the workstream — data analysis, page variants, screens, content — while your team continues the rest. The studio does not need to “take over the account” to be useful.

If the workstream is already well framed by an agency or an in-house team, we fit their rituals. The call is there to see where we plug in, not to rebuild everything.

SYS/CALL

Next step

A 30-minute call to talk about a conversion funnel, a product to build or data to analyse.

SYS/VIEW