Skip to content

Rishi Rai

All Writings

Combining Software Engineering Discipline with AI Application Development

Build deterministic boundaries, versioned data and prompts, layered tests, observability, secure tool use, and reversible delivery around AI.

Engineering analysis

For: Software engineers applying established engineering practices to probabilistic AI application components.

2026-10-02

Software Engineering for AI Applications

AI applications still need the disciplines used to build dependable distributed systems, but their boundaries are less obvious. Model behavior is probabilistic, prompts act like executable configuration, retrieved data changes over time, and providers can alter performance without a code deployment. The engineering response is not to pretend that generation is deterministic. It is to isolate uncertainty behind explicit contracts and make the surrounding system testable, observable, and reversible.

A useful design separates decisions that software can enforce from judgments delegated to a model. Authentication, authorization, money movement, state transitions, and destructive actions belong in deterministic code. Models can propose classifications, drafts, or plans, but trusted components validate and execute them.

Architecture: A Deterministic Core Around a Probabilistic Component

Structure the system as a pipeline with replaceable modules: input normalization, policy enforcement, context assembly, model invocation, output validation, action execution, and audit recording. The model adapter should expose a provider-neutral interface while retaining provider metadata for diagnosis. Business workflows should depend on an application contract, not a specific model name or response shape.

  1. Boundary layer: authenticates callers, validates request syntax, and assigns an idempotency key.
  2. Deterministic core: loads authorized data, applies business rules, and chooses an approved workflow.
  3. AI module: produces a typed proposal using supplied context and constrained tools.
  4. Validation layer: checks schema, policy, evidence, and domain invariants.
  5. Execution layer: performs permitted effects and records an auditable result.

The model may recommend an action. The application remains responsible for deciding whether that action is valid and allowed.

Define Contracts Before Prompts

Specify inputs, outputs, failure modes, and side effects independently of prompt wording. An output contract should distinguish required fields, optional evidence, confidence or uncertainty, and refusal states. Prefer enumerated values and bounded collections over prose that downstream code must parse. Validate responses with the same rigor as an external API.

type ReviewProposal = {
  decision: "approve" | "reject" | "needs_human"
  reasons: string[]
  evidenceIds: string[]
  proposedActions: ProposedAction[]
  contractVersion: "3"
}

proposal = model.generate(context, schema)
validated = validateSchema(proposal)
authorized = validateActions(validated.proposedActions, actor)

if (!authorized.ok) return queueHumanReview(proposal)
return executeIdempotently(authorized.actions)

Version the contract and reject unknown versions. A schema-valid response can still violate domain rules, so validate relationships such as totals, ownership, time ranges, and permitted transitions. Never allow free-form model text to become a database query, shell command, or privileged tool argument without strict parsing and authorization.

Create Modular Boundaries

Keep prompt templates, retrieval, model adapters, tools, validators, and workflow orchestration in separate modules. This permits independent tests and controlled replacement. The prompt builder should accept typed context and return a message structure. The retrieval module should return passages with source identifiers and data versions. The model adapter should handle timeouts, usage metadata, and provider errors. The validator should know the contract but not how the model was called.

Avoid a universal AI service that accepts arbitrary prompts and tools. It becomes difficult to authorize, evaluate, or maintain. Prefer narrow capabilities such as summarizeIncident, extractInvoiceFields, or draftReply, each with a bounded contract and evaluation set. Shared infrastructure can remain reusable beneath those interfaces.

Keep State and Data Reproducible

Store immutable identifiers for the model configuration, prompt template, output contract, retrieval index, embedding model, tool definitions, and policy rules used by each request. Record source document identifiers and revisions rather than copying sensitive documents into logs. This provenance makes failures reproducible and supports targeted reprocessing when a component changes.

Data versioning matters because the same code and prompt can produce a different answer after an index refresh. Build indexes through repeatable jobs, publish them under a version, validate them, then atomically promote an alias. Retain the previous compatible version for rollback. Apply migrations to stored conversations and model outputs as carefully as database migrations; readers must handle old versions until backfills complete.

Test in Layers

Unit tests should cover deterministic logic: prompt assembly, token budgeting, authorization filters, schema parsing, fallback selection, and tool argument validation. Use recorded provider responses at module boundaries so most tests are fast and stable. Contract tests should run against each supported provider or model configuration to detect changes in structured output and tool behavior.

Evaluation tests address behavior that ordinary assertions cannot fully capture. Maintain a versioned dataset representing common tasks, edge cases, adversarial input, ambiguous requests, refusals, and historically observed failures. Combine exact checks with task-specific rubrics and periodic human scoring. Evaluation results should include slices by language, customer type, document length, and risk level so aggregate scores do not hide regressions.

  • Test malformed and partially valid model responses.
  • Inject timeouts, rate limits, retrieval gaps, and unavailable tools.
  • Verify that duplicate requests do not repeat side effects.
  • Confirm that unauthorized evidence never enters model context.
  • Run load tests with realistic token distributions, not uniform short prompts.

Build Observability Around Decisions

Trace the complete request across retrieval, generation, validation, tools, and persistence. Record durations, token counts, cost estimates, cache outcomes, retries, model configuration versions, validation failures, and human review outcomes. Use redaction and sampling so observability does not become a second sensitive-data store.

Operational metrics need behavioral companions. Track schema acceptance, grounded evidence coverage, tool success, refusal rate, repair attempts, escalation rate, and user corrections. Alerts should focus on changes relative to a stable baseline and should be segmented by workflow. A global success rate can remain flat while one critical use case fails.

Deploy and Roll Back Safely

Treat prompts, models, retrieval settings, and policy rules as deployable artifacts reviewed alongside code. Pin versions where the provider permits it and maintain a manifest that identifies the complete runtime configuration. Before promotion, run offline evaluations and contract tests. Then use shadow traffic or a small canary cohort with explicit stop conditions.

release = {
  app: "2026.10.2",
  prompt: "incident-summary-7",
  contract: "review-proposal-3",
  modelPolicy: "router-12",
  retrievalIndex: "runbooks-41"
}

if canary.schemaFailureRate > baseline.limit:
    rollback(release.previousCompatible)
if canary.humanEscalationDelta > policy.limit:
    haltPromotion()

Rollback must restore a compatible set, not just the previous prompt. A prompt may depend on a new schema or tool definition. Keep prior manifests and verify rollback paths in staging. For long-running jobs, store the chosen manifest with the job so a mid-deployment change does not mix configurations.

Design Human Review as a Workflow

Human review is appropriate for high-impact actions, weak evidence, novel cases, and repeated validation failure. Present reviewers with the proposed output, relevant source excerpts, triggered rules, and a clear set of actions. Do not force them to reconstruct context from raw traces. Capture structured reasons for edits or rejection so recurring problems can become tests.

Set service levels and ownership for review queues. A review requirement without staffing becomes silent failure. Guard against automation bias by showing uncertainty and evidence before the recommendation when practical, and periodically measure agreement between reviewers.

Security and Abuse Resistance

Assume user text, retrieved documents, and tool responses can contain prompt injection. Mark them as data, minimize their authority, and enforce tool permissions in code. Use least-privilege credentials per tool and tenant, short-lived tokens, network restrictions, and explicit allowlists for destinations. Validate every tool call against the authenticated actor and current resource state.

Classify data before sending it to a provider. Redact secrets and unnecessary personal data, define retention rules, and verify regional and contractual requirements. Protect training and evaluation datasets from poisoning through provenance, access control, review, and anomaly checks. Rate limits, quotas, bounded loops, and token ceilings constrain denial-of-wallet attacks.

Trade-offs and Common Mistakes

  • Over-abstraction: a provider-neutral adapter is useful, but hiding every capability can prevent use of valuable model-specific features. Keep a stable core and expose optional capabilities explicitly.
  • Snapshot tests alone: exact text snapshots are brittle and miss semantic defects. Assert contracts and invariants, then evaluate meaning separately.
  • Silent fallback: changing models may alter safety, context limits, or tool support. Log the fallback and verify compatibility first.
  • Unbounded agents: open-ended loops create unpredictable cost and side effects. Limit steps, time, tools, and cumulative spend.
  • Prompt-only policy: instructions can guide behavior but cannot replace authorization or transaction rules.
  • Missing ownership: every prompt, evaluation set, data source, and review queue needs a maintainer and retirement path.

Maintainability as a Product Requirement

Document each capability's contract, data dependencies, quality thresholds, operational limits, and failure behavior. Remove obsolete prompts and model configurations instead of accumulating hidden fallbacks. Schedule evaluation refreshes as workflows and data evolve. Prefer small, observable pipelines over elaborate autonomous designs when both meet the requirement.

Conclusion

Reliable AI applications emerge from clear contracts and controlled uncertainty. A deterministic core enforces policy and state, modular AI components produce bounded proposals, and validators protect every side effect. Versioned data, layered tests, decision-focused observability, compatible rollback, meaningful human review, and least-privilege security make the system operable over time. These practices let teams change models and prompts without surrendering the engineering properties that production software requires.