Governing Skills in the Enterprise 101

How to protect your knowledge and IP without drowning in hosting and governance overhead.


Most AI rollouts still look the same. You buy licenses, point people at a chat window, and hope for the best. Every prompt starts from zero, and every output gets reviewed by hand, forever.

Most people are prompting and calling that AI usage. But usage alone doesn’t compound: for a repeatable task, that improvement only comes from turning the prompt into a skill, one governed by IT the same way any other enterprise system is, not left sitting in someone’s chat history.

Core principle: slope beats y-intercept

Most people size up an AI system by its first draft: is this answer good enough to ship? That’s a y-intercept question, and it’s the wrong one. What actually determines the return is the slope: does the tenth run need less correction than the first, and does the hundredth need almost none?

We mapped the mechanics of that pipeline in “The Future of Work.” Compounding only starts once each stage’s standard is codified into a skill, not left to reset every time the chat window closes. Skip that, and you’re paying the same correction cost on run 100 as run 1, no matter how many people are prompting.

Skill Management 101

A skill file by itself is just better documentation. What actually keeps quality climbing and review time dropping are three plumbing problems that have nothing to do with any single skill’s content, and every firm running AI at real scale needs all three solved the same way.

  • Human feedback channels. Every edit a reviewer makes has to land back inside the skill, not just in that day’s output. That correction is the raw material the whole loop runs on.
  • Evals. Structured tests that check a skill’s output against a standard, run before a person ever sees the draft. Evals are also what let a team have an honest conversation about which model a task actually needs: whether a cheaper model already clears the bar for routine work, or whether one specific step still needs the frontier model.
  • Cost and frequency, tracked and reported up. Someone at the firm needs a real number for leadership: what a given skill costs to run, how often it runs, and what outcome it’s tied to. Without that, AI spending reads as a vague IT line item instead of a process with a return.
Skill runs
Produces a draft against the current version of the skill.
Reviewer edits or approves
A human makes the correction the skill didn't yet know to make.
Edit + eval data feed back in
The correction and its eval result get incorporated into the skill itself.
Next run starts closer to correct
The plumbing underneath every skill, regardless of what the skill itself does.

Democratizing tacit knowledge (but not too much)

The most common objection to partnering with a technology vendor on any of this is fair: “I don’t want to hand over my playbook, the standard operating procedures that took years to refine, to whoever’s running the model underneath.”

That worry has gotten louder, not quieter. Microsoft CEO Satya Nadella raised it directly in a widely shared essay this past June:

“The last thing any of us want is a world where every company across every sector is ceding value to a few models that eat everything they see.”

“A small number of AI systems capturing all the economic returns, while entire industries find their knowledge commoditized right out from underneath them.”

— Satya Nadella, Microsoft CEO (source)

His fix is two kinds of capital working together: human judgment, and what he calls “token capital,” the AI capability a firm builds and owns. That token capital is just the human-feedback-and-eval loop covered above, encoded so an entire firm, not just your best few people, can execute your standards, keep improving, and run at LLM cost instead of headcount cost, all without resetting when you swap models.

Sinusoidal’s philosophy: govern what’s common, protect what’s yours

Governing skills well means splitting what should be common infrastructure from what should never leave your walls. Where and how that infrastructure runs, your own cloud or a managed one, model-agnostic by design, built to your compliance requirements, is its own question we’ve written up in “Managing Execution Environments and Data Connectors.” Here’s how we split the responsibility itself:

Common, run for everyoneOurs, opt out anytimeAlways yours, we never train on it
Cloud orchestration, your infrastructure or ours, based on your compliance needsStarter skill templates plus built-in improvement loopsYour own skills
A model-agnostic control planeLearning from edits made to our own templatesYour past work and data
Skill lifecycle management and versioningOur caching layer to cut repeat costsYour own learning loop
SLAs and spend reporting, cost-per-skill includedYour custom triggers

A few things worth pulling out of that table on their own:

  • We don’t train our own models on your private work, full stop, regardless of which tier you’re using.
  • Bring your own skills and keep them yours. We don’t need to see inside them to run them.
  • Model-agnostic means no lock-in to one lab’s weights, so switching providers doesn’t mean rebuilding your governance from scratch.

Civil engineering firms feel this directly: a QA/QC checklist tuned to an agency’s past comments, a proposal voice that’s already won work with a client, a way of describing stormwater risk that’s distinctly your firm’s. That’s the tacit knowledge worth protecting, exactly the “alpha” Nadella’s warning is about.

It shouldn’t get flattened into a generic template or handed to whoever’s running the model underneath. We built Sinusoidal to encode your SOPs, checklists, and proposal voice as governed skills that stay yours, versioned and improvable, without giving up what makes your RFPs and plan reviews different from a competitor’s.

If your firm is thinking through what that looks like, we’d welcome the conversation.

Talk to Sinusoidal