← CNSLT JOURNAL
SAFE AI / BUILD IN PUBLIC

How Do We Build Secure and Trustworthy AI? Introducing SAFE AI

I started exploring AI through prompts. Now I want to understand what it takes to build AI systems that are useful, secure, and accountable—not just impressive in a demo.

CNSLT ACADEMYOCTOBER 11, 20267 MIN READ

IN BRIEF

Building secure and trustworthy AI means combining AI engineering with security testing, technical guardrails, risk management, and human oversight. SAFE AI is CNSLT Academy's build-in-public capstone exploring that lifecycle through practical experiments.

In mid-2023, when I was first experimenting with ChatGPT, I remember asking it for a very long list of websites. Getting hundreds of items in one response was harder than I expected: outputs could stop early, drift, or fail to meet the format I had requested. It was a small interaction, but it captured how new and unpredictable these tools still felt.

Since then, the way people use AI has expanded. Modern systems can be connected to web search, code execution, document stores, APIs, and other tools. Depending on their setup and permissions, they can carry out parts of a workflow rather than only return text. That is a meaningful shift—and it changes the security questions we need to ask.

The story is not that AI is inevitably going to become a rogue decision-maker. That is a possible scenario people debate, not a conclusion we can treat as inevitable. The more immediate challenge is that real systems can already be wrong, manipulated, over-permissioned, poorly evaluated, or deployed without adequate oversight.

The question is no longer only “What can AI do?”

It is also: what can this system access, what actions can it take, what happens when it is manipulated, and who is accountable when it fails?

A model can produce a convincing but false answer. A document retrieved by an AI application can contain hostile instructions. A connected tool can have permissions that are broader than the task requires. A governance policy can look complete on paper while nobody tests whether the technical controls actually enforce it.

These are not problems solved by writing a better prompt alone. They require a combination of engineering, evaluation, security testing, risk management, and operational controls.

Why SAFE AI?

I am building SAFE AI as the flagship capstone for CNSLT Academy. The aim is to learn the full lifecycle by working on one evolving project—not to collect disconnected tutorials or pretend that reading about security is the same as practising it.

Understand AI. Build AI. Break AI. Secure AI. Govern AI.

That sequence is the project’s working principle. Every stage should leave behind something observable: an experiment, a working feature, a reproducible test, a mitigation, or documented control evidence.

1. Understand

Learn the fundamentals: tokens, embeddings, context windows, prompting, retrieval, model limitations, and evaluation. The goal is to know what the system does and where uncertainty enters—not to memorize terminology.

2. Build

Create a practical AI application with document retrieval, structured outputs, logging, and carefully scoped tool access. Start small enough to understand the data flow and the trust boundaries.

3. Break

Test the application in a controlled environment. Try prompt injection, jailbreaks, retrieval poisoning, data leakage, and attempts to manipulate outputs or misuse tools. Record the exact conditions, impact, and reproducible steps for each failure.

4. Secure

Add and test mitigations: least-privilege permissions, input and output validation, separation of trusted instructions from untrusted content, secrets protection, rate limits, and human approval for consequential actions. A control is not proven merely because it was added; it needs retesting.

5. Govern

Connect the technical work to policy, risk assessment, accountability, monitoring, and evidence. Map the application to relevant guidance such as the NIST AI Risk Management Framework and the OWASP guidance for generative AI applications.

The governance layer must do more than generate advice

One part of the capstone will examine AI use cases from an information-security and governance perspective. The system should help assemble evidence, identify risks, map relevant controls, and recommend a decision path. For a proposed AI use case, that path could be:

The model should not be the final authority on its own. Deterministic policy checks, evidence requirements, permission enforcement, and human review need to sit around it. The project will explore where AI can assist a decision and where the system must enforce a rule independently.

What I will publish as I build

This is a learning journey, so I will document the work rather than present a finished product before it exists. That means publishing:

Over time, the most useful material can become structured lessons, labs, templates, and a course on CNSLT.IN. The course should be shaped by what the project actually demonstrates, not by promises made in advance.

The goal is practical competence, not fear

We do not need to predict a single dramatic future to take AI security seriously. We need to understand systems as they are deployed, identify where they can fail, reduce avoidable risk, and make accountability concrete.

SAFE AI is my attempt to learn that discipline in public—and to build a capstone that connects AI engineering, red teaming, security controls, and governance in one place.

The first version will be imperfect. The tests should be honest. The fixes should be measurable. And the lessons should be useful to someone building their own skills.

Frequently asked questions

What is AI security?

AI security is the practice of identifying and reducing risks in AI models and applications, including prompt injection, sensitive-data exposure, unsafe tool use, and failures in the systems around a model.

What is AI governance?

AI governance is the set of policies, roles, risk processes, controls, monitoring, and evidence used to guide how an organization develops, approves, deploys, and oversees AI systems.

How do you test an AI application for security?

Start with a threat model and defined scope, then test realistic failure modes such as prompt injection, retrieval poisoning, data leakage, excessive permissions, and unsafe outputs. Record reproducible evidence, apply mitigations, and retest.

Can AI security and governance be combined?

Yes. Technical tests identify how a system can fail; governance processes determine acceptable risk, required controls, accountability, and when human approval is needed. The two are strongest when linked to the same system and evidence.

What is SAFE AI?

SAFE AI is CNSLT Academy's planned practical capstone for learning AI engineering, red teaming, security controls, and governance through one evolving project. It is a work in progress, not a claim that a finished product has already been validated.

FOLLOW THE BUILD

Follow the CNSLT journal for the experiments, build notes, red-team tests, and governance work behind SAFE AI.

EXPLORE THE JOURNAL ↗