Signpost and Responsible AI
Why Generative AI in support of Signpost programming goals?
Signpost supports responsive information programs using a common digital architecture that supports the creation of original web-based content and uses a common backend for connecting digital communications channels. Since 2015, Signpost has launched programs in over 30 countries and 25 languages, and has registered approximately 20 million users of its information products, and offered roughly 500,000 one-to-one information counselling sessions.
Our one-to-one communication has always been a hallmark of Signpost's program strategy, opening doors for personalized support to people who would like extra help. This has always been limited by staff time. This philosophy was further legitimized by a Randomized Controlled Trial (RCT) in Greece, which identified that information sessions in the form of 1-on-1 communication increased the uptake for humanitarian services and increased the overall trust in institutions.
The scale of that staff-time constraint became concrete in practice. When Signpost launched in Afghanistan in 2021, a minimal-cost social media campaign in late 2023 rapidly generated 48,000 in-bound inquiries and support requests within weeks — a volume that overwhelmed a team of three personnel, forcing them to take the post down and stop promoting two-way information counseling due to lack of staff capacity. That's the gap generative AI was meant to close.
The two of these — architecture & data, plus a high-value, evidenced use case — constitute the business case for leveraging AI. The hypothesis: if we can leverage AI to support one-to-one communications with Signpost users without sacrificing quality and while ensuring safety and privacy of our clients, we will be able to safely expand our one-to-one services at dramatically lower cost and thereby increase the impact of Signpost programs.
Early Pilots
Starting with the Google Impact Accelerator, we built a RAG-based agent — Signpost AI Information Assistant — that used Signpost articles as its knowledge base, with prompt design intended to keep outputs in line with our quality and safety standards. This RAG approach let the LLM draw on verified external knowledge rather than relying on its own static training data, addressing the tendency of conventional LLMs to fabricate responses. The first version of the Vector Database contained almost 30,000 articles created by Signpost to meet users' self-expressed information needs.
The tool was piloted for six months, with deployment starting in September 2024 in Greece and Italy, and in October 2024 in Kenya and El Salvador. Each country worked from a distinct, contextualized knowledge base and localized system prompts. The pilot was structured in four phases — deployment, testing and refinement, impact assessment, and scaling/sustainability planning — and was run with a human in the loop throughout: moderators reviewed, evaluated, and scored every AI-generated response before it reached a client.
The results were strong, though not unambiguous. Overall performance improved over the course of the pilot, rising from 51.68% to 76.81% across all countries. By country, Greece processed 885 inquiries with 62% of AI-generated responses rated "good"; Italy processed 323 inquiries with 67% rated good; Kenya tested 719 inquiries with 311 rated good or usable; and El Salvador processed 241 inquiries with 70% rated good or usable with minor edits.
Staff experience with the tool was consistently positive. Moderators reported that without the chatbot, writing a response to a complex user question took between 7 and 25 minutes — with AI assistance, just 2 to 5 minutes. 87.5% of moderators in Greece and Italy, and 80% of moderators in Kenya, said they'd recommend the program continue using the AI after the pilot ended, and every moderator who participated in the pilot requested to keep using it afterward.
But the pilot was also clear-eyed about the tool's limits. Given the high-stakes humanitarian context of Signpost's work, the Information Assistant appeared safe only with a human in the loop: roughly 85% of responses were judged safe, about 79% fully met the bar for trauma-informed language, and 75% fully met the bar for a high-quality, client-centered response — meaning the remaining share could have introduced harm to a client without human review. Hallucinations occurred in roughly 13% of responses in both Greece and Italy, requiring significant moderation. The clearest driver of quality, across every country, was the knowledge base itself: the tool performed better wherever its knowledge base more fully reflected the range of user needs, and suffered where content was thin or missing context.
So rather than unlocking a jump to a human-on-the-loop model, the pilot's real finding was more foundational: context — a robust, context-specific knowledge base — is the single most critical component of an AI system like this one, and the pilot's results offered a blueprint for the AI systems that would come next in humanitarian information delivery.
Agent Building
We created an agent-building platform, originally called "Signpost AI," to meet the much more complex requirements of building agents like this responsibly and at scale. The platform's technical architecture is built to give teams an interface to create and configure agents for highly customizable workflows.
Signpost AI was designed specifically to create, test, and safeguard AI agents for humanitarian use with tighter safety guardrails than a generic chatbot — and building it that way unlocked the ability to create agents across sectors, from protection and education to refugee resettlement and anti–human trafficking work. Rather than relying on generic chatbots, IRC designs its own purpose-built agents — a process closer to writing a job description than configuring off-the-shelf software: defining roles, responsibilities, safety protocols and guardrails, workflows, and knowledge boundaries for each one.
Signpost is piloting its first human-on-the-loop agent in Mexico at the time of this publication to determine if greater efficiency gains are possible from this technical approach without sacrificing safety and quality standards.
Transformation - from Signpost to IRC’s new Technology for Programs Delivery Team
With a growing portfolio of support across programs in different sectors and technical units at the IRC, and an ever-increasing demand the Signpost engineers, we decided that only through the centralization of the technical team could we meet the highest priority work for the IRC. This team now officially serves the greater organization as the applied technology for program team and still supports programs like Signpost responsive information services.
What’s next for Signpost?
Signpost is now located in the IRC’s Violence Prevention & Response Unit (VPRU) and will continue its mission to ensure high quality responsive information services in humanitarian crises worldwide.