Guided private-beta package
AI Software Teams. Controlled.
Evaluate how ForgeLayer connects prompts, agents, code review, verification, approval, Work Records, Ledger context, and Analytics in one controlled journey.
Package summary
What this guided evaluation demonstrates
What teams can evaluate
Connected product proof, module by module
Workspace
Guided beta- Launch the guided private-beta Control Flow from a single operational workspace.
- Surface AI work requiring approval, changes, verification, or blocking decisions.
- Inspect recent safe Review Run summaries and derived Ledger activity.
- Show deterministic AgentOS routing and supporting-agent context.
PromptForge
Guided beta- Prompt quality scoring
- Prompt security scoring
- Injection-risk detection
- System-prompt leakage detection
AgentOS
Guided beta- Agent registry
- Task routing
- Capability matching
- Control-risk comparison
Review Runs
Guided beta- Accept safe review intake from GitHub webhooks, imported diffs, and deterministic demos.
- Collect report, repository-policy, check-preview, comment-preview, and warning evidence.
- Evaluate RepoBrain policy context and protected or sensitive path matches.
- Classify control risk without treating a recommendation as human approval.
MergeGuard
Guided beta- PR diff inspection
- RepoBrain policy context
- Risk classification
- Review Run lifecycle
AI Work Ledger
Guided beta- Connect prompt, routing, review, approval, verification, and outcome records.
- Separate observed facts from recommendations, draft decisions, and human approvals.
- Link evidence and related records without exposing raw prompts, diffs, secrets, or payloads.
- Surface approval and verification state before work is trusted or merged.
AI Work Analytics
Guided beta- Summarize deterministic AI work volume and source-module distribution.
- Compare recommended primary and supporting agent routing without scoring agent quality.
- Show capability-fit signals and control risk as required oversight.
- Explain review, verification, approval, policy, block, and changes-requested outcomes.
Supported pilot scenarios
Choose one bounded task
Secure a payment webhook
Review and harden payment webhook signature verification, replay protection, tests, and rollback evidence before merge.
Expected outcomeShow how sensitive payment and session changes remain blocked behind evidence, policy, and explicit human approval.
9 connected modulesHarden a risky prompt
Review a support-agent prompt for injection, secret leakage, unsafe tools, data export, and missing human approval gates.
Expected outcomeShow how unsafe prompt authority, injection, leakage, and tool risk become visible before an agent receives the instruction.
8 connected modulesInvestigate failing tests
Investigate a failing checkout test suite, separate observed failures from proposed fixes, and require verification before code changes are trusted.
Expected outcomeShow how observed failures stay separate from proposed fixes until verification evidence and a human decision are present.
8 connected modulesGuided product walkthrough
Workspace to readiness, with control at every stage
Workspace
Frame the task, attention queue, and current boundaries.
- Product proof
- One control center connects the active deterministic task to the product modules.
- Control boundary
- Demo workspace only; no production identity or persistence.
Pilot
Select one repeatable scenario and evaluation objective.
- Product proof
- Three stable scenarios exercise different risk and review paths.
- Control boundary
- Local scenario selection is non-persistent.
Journey
Trace the task across the complete control lifecycle.
- Product proof
- Module ownership, gates, and next actions remain connected.
- Control boundary
- Recommendations and previews only.
PromptForge
Evaluate prompt quality and security before agent use.
- Product proof
- Deterministic scoring, rewrite, and security tests are inspectable.
- Control boundary
- No model execution or raw prompt persistence.
AgentOS
Route work to suitable agent profiles under oversight.
- Product proof
- Primary and supporting recommendations explain fit and control risk.
- Control boundary
- No agent is contacted, executed, or granted access.
Control Flow
Connect routing, review, verification, and approval gates.
- Product proof
- Blocked actions and next controlled steps stay visible.
- Control boundary
- Draft plan only; no orchestration.
Review Run
Inspect policy context, risk, evidence, and review status.
- Product proof
- Safe simulated Review Runs preserve evidence and decision context.
- Control boundary
- Demo data is simulated and non-persistent.
MergeGuard
Translate review evidence into a controlled merge recommendation.
- Product proof
- Check and comment previews remain distinct from human approval.
- Control boundary
- No GitHub posting, repository write, or automatic merge.
Verification
Separate available, missing, failed, and simulated evidence.
- Product proof
- Verification state and remediation remain explicit.
- Control boundary
- Some evidence is simulated and cannot establish a production pass.
Human Approval
Require a person to approve, request changes, or block.
- Product proof
- The system recommendation never substitutes for a human decision.
- Control boundary
- Approval is local-only and non-persistent.
Work Record
Record scope, routing, risks, evidence, and decisions.
- Product proof
- Connected record details preserve the controlled-work narrative.
- Control boundary
- Draft preview records are not production records.
Ledger
Connect tasks, decisions, evidence, and outcomes.
- Product proof
- Linked deterministic records provide traceability in preview.
- Control boundary
- No production persistence, retention, or immutable audit claim.
Analytics
Summarize control flow, oversight, evidence, and blockers.
- Product proof
- Management signals derive from seeded ForgeLayer records.
- Control boundary
- No production telemetry, customer benchmark, or performance guarantee.
Evaluation
Assess evidence and control coverage qualitatively.
- Product proof
- Demonstrated, partial, future, and blocked states stay distinct.
- Control boundary
- No customer score or production-readiness claim.
Findings
Convert observed strengths and gaps into follow-up actions.
- Product proof
- Findings connect to evidence, owners, blockers, and completion criteria.
- Control boundary
- Local actions do not resolve production prerequisites.
Readiness
Confirm the guided-beta checklist and production blockers.
- Product proof
- Guided-beta readiness remains separate from self-service deployment.
- Control boundary
- Production auth, persistence, billing, and orchestration are not live.
Evidence and control coverage
What reviewers can inspect and decide
Task scope
- app/api/payments/webhook.ts
- tests/payment-webhook.test.ts
Scenario evidence
- Webhook test plan drafted
- Protected payment path matched
- Security review required
Review record
- Secure payment webhook replay handling (demo-task-payment-webhook)
Ledger evidence
- 28 linked deterministic evidence labels
Analytics evidence
- 12 fixed demo activity records
- 53 linked evidence labels
Pilot evaluation criteria
Evaluate product proof qualitatively
Task understanding
Supporting module: Workspace
Inspect supporting surfaceAgent fit
Supporting module: AgentOS
Inspect supporting surfaceControl-risk visibility
Supporting module: Control Flow
Inspect supporting surfaceReview evidence
Supporting module: Review Runs
Inspect supporting surfaceHuman approval
Supporting module: Pilot
Inspect supporting surfaceTraceability
Supporting module: Journey
Inspect supporting surfaceLedger completeness
Supporting module: AI Work Ledger
Inspect supporting surfaceAnalytics visibility
Supporting module: AI Work Analytics
Inspect supporting surfaceSafety-boundary clarity
Supporting module: Readiness
Inspect supporting surfaceProduction readiness
Supporting module: Production Architecture
Inspect supporting surfaceFindings and action-plan preview
Turn the guided evaluation into a bounded follow-up
Review the findings with the pilot team
Record local owner and status choices, then export the discussion manually if needed.
3 supporting findingsDefine pilot evidence and repository-policy ownership
Compare the agreed requirements with Review Runs, RepoBrain, and MergeGuard previews.
2 supporting findingsImplement identity and workspace isolation before durable expansion
Use the production architecture plan to sequence identity, RLS, and durable records.
3 supporting findingsDesign side-effect gates before considering execution
Keep all execution flags disabled and review the GitHub Checks and production architecture plans.
2 supporting findingsPrivate-beta readiness checklist
Ready for guidance is not ready for deployment
Product experience
Ready for guided beta- Workspace, Pilot, Journey, and module surfaces are connected.
- Three deterministic pilot scenarios are available.
- The private-beta lifecycle is understandable end to end.
Safety boundaries
Ready for guided beta- Execution, repository writes, GitHub posting, and auto-merge remain disabled.
- Human approval remains explicit.
- Preview and simulated states remain labeled.
Pilot scenario readiness
Ready for guided beta- Secure a payment webhook: deterministic scenario available.
- Harden a risky prompt: deterministic scenario available.
- Investigate failing tests: deterministic scenario available.
Evaluation readiness
Ready for guided beta- Qualitative evaluation categories exist.
- Findings and action-plan previews are connected.
- No customer score or production metric is presented.
Evidence and traceability
Partially ready- Safe review, policy, verification, Work Record, and Ledger evidence is available.
- Some verification evidence remains simulated.
- Production retention and immutable audit history are not live.
Documentation
Ready for guided beta- Open /api/health and confirm the deployed build marker is current.
- Open /readiness and confirm ForgeLayer is private-beta demo-ready, not production self-serve.
- Open /demo and follow the 10-15 minute reviewer walkthrough.
- Open /onboarding and confirm the safe pilot path is clear.
Team onboarding
Partially ready- A guided session owner and human approver are required.
- Self-service onboarding is not live.
- The team must provide feedback on workflow fit and missing controls.
Production prerequisites
Future production requirement- Production workspace identity, membership, and role enforcement
- Workspace-scoped persistence with reviewed RLS policies
- Durable approval, evidence, Work Record, Ledger, and audit event storage
- Explicit repository and GitHub App permission boundaries
- Idempotent workers, retry policy, monitoring, and kill switches
- Controlled execution policy with scoped agent and tool permissions
- Production analytics consent, retention, privacy, and telemetry boundaries
Allowed checklist states: Ready For Guided Beta · Partially Ready · Blocked By Design · Future Production Requirement
Pilot engagement outline
A focused technical evaluation, not a sales commitment
- 01Select one scenario
Choose the payment webhook, risky prompt, or failing-test journey.
- 02Walk through the controlled journey
Use Workspace, Pilot, and Journey to frame module ownership and gates.
- 03Inspect routing and governance
Review agent fit, supporting roles, authority boundaries, and required oversight.
- 04Inspect review and verification evidence
Separate observed evidence, missing evidence, and simulated verification.
- 05Make local approval decisions
Approve, request changes, or block without triggering any external action.
- 06Inspect Work Record and Ledger
Trace scope, evidence, risk, approval, and linked record context.
- 07Review Analytics
Discuss oversight, blockers, evidence coverage, and workflow-state signals.
- 08Complete Evaluation and Findings
Capture demonstrated value, unresolved gaps, and production prerequisites.
- 09Agree on next steps
Choose a bounded follow-up without pricing, contracts, deadlines, or commitments.
Team participation
- Choose one bounded pilot scenario and identify the engineering context.
- Provide a technical lead or reviewer for the guided evaluation.
- Inspect evidence and make explicit local approval decisions.
- Identify missing controls, confusing concepts, and workflow-fit concerns.
- Confirm which production prerequisites would be required before real use.
ForgeLayer participation
- Guide the team through the deterministic control journey.
- Explain every recommendation, preview state, and safety boundary.
- Keep execution, writes, posting, merge, and persistence disabled.
- Capture findings and next actions in preview without external sending.
- Separate demonstrated value from production prerequisites and future work.
Boundaries and blockers
What remains outside this package
Workspace identity and RLS are prerequisites for durable records
Owner: Platform Owner
- Production auth not live
- Workspace RLS not implemented
Repository writes and automatic actions remain blocked by design
Owner: Security Reviewer
- No controlled execution boundary
- No production write authorization