A design paper, before a deployment
A paper published October 1 in Ethics and Information Technology proposes boundaries for religious-learning chatbots, including a warning against giving them names that suggest clerical authority. Christos Papakostas, affiliated with the National and Kapodistrian University of Athens, presents a conceptual framework rather than an operating product. No datasets were generated or analysed. Source: link.springer.com
The proposed system combines an approved theological collection, constrained retrieval and response templates, Moodle integration, citations and human oversight. Its main setting is upper-secondary or higher-education religious learning. The author excludes sacramental confession, clinical counselling and autonomous pastoral care from its intended role.
The paper even questions its own provisional name, “Ask Father AI,” recommending a learning-assistant name for any implementation. It calls for empirical work before classroom use. That qualification governs how the proposal should be read: descriptions of safeguards are design requirements, not results demonstrating that a chatbot has met them.
For schools and churches considering AI, the development makes a familiar ethical aspiration more concrete. It also exposes the distance between writing down a boundary and showing that a system reliably respects it.
A citation is an invitation to examine the answer
One relevant technical precedent is ALCE, the citation-evaluation benchmark introduced by Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen in 2023. Its public research repository separates fluency, correctness and citation quality, using three question-answering datasets. It also makes evaluation code and human-evaluation material available. Source: github.com
The distinction matters more here than any historical model ranking. These are separate properties to examine. A readable answer can be wrong. A correct answer can fail to show its basis. A source reference can be present without supporting every assertion attached to it. This is a reason to test each relationship, rather than count hyperlinks.
ALCE was not a test of the new pastoral framework or a religious classroom. Its relevance is methodological. A religious-learning evaluation would additionally need qualified judgment about whether a source is represented in context and whether the answer identifies the interpretive tradition it is using. Those are proposed adaptations, not results from the benchmark.
The institution must make the boundary real
There is a longer ethical argument behind the gap between principles and practice. In a 2019 paper, Brent Mittelstadt warned that agreement on high-level values does not establish agreement about how to implement them. His analysis calls for visible accountability, review procedures, documentation and attention to organizational incentives. It also argues against reducing difficult ethical choices to technical design alone. Source: arxiv.org
The article predates the present classroom proposal. It is used here as a conceptual comparison, not as a statement about today's laws or a review of Papakostas's work. Its question remains useful: what happens when an appealing principle becomes inconvenient?
Consider a hypothetical course in which a disputed answer has already been shared with an entire class. Someone must have authority to pause the system, correct the teaching and explain the change. A committee's name on a diagram does not by itself establish who can act before its next meeting. The operating arrangements are part of the ethical design.
Educational value requires its own evidence
UNESCO's published introduction to its 2023 guidance on generative AI in education calls for human-centred, age-appropriate ethical validation and pedagogical design. The document's aims include meaningful, equitable use, rather than adoption for its own sake. This is international guidance, not certification of a particular product. Source: unesco.org
Applied to religious learning, that makes the central educational question more demanding than whether students receive an immediate answer. Can they explain the source afterward? Can they recognize disagreement? Do they become better able to ask a teacher a precise question? These are suggested evaluation questions, not outcomes demonstrated by the new paper.
There is a reasonable positive case for assistance. A carefully bounded tool might help a learner find a passage, revisit a class topic or prepare for discussion when a teacher is unavailable. That possibility deserves testing. It should not be confused with proof that learners understand more, trust appropriately or receive better pastoral attention. A useful research program would measure the intended benefit and the possibility of unintended dependence.
Christian perspective: equip people to participate
Ephesians 4:11–16 describes pastors and teachers equipping the church for service and maturity. The passage moves from differing gifts toward a community that grows together in truth and love under Christ. It does not define educational software, and Christians differ over how its ministry language maps onto church offices. Ephesians 4:11-16 (NIV)
Its practical challenge for a Christian institution is nevertheless specific: the purpose of teaching is to form people capable of faithful participation. A machine's ability to produce a polished explanation is therefore an incomplete measure of success. The learner's capacity to understand, question, serve and join communal life matters too.
This need not exclude technology. A searchable collection or a carefully checked explanation can help a person enter a conversation. But a tool presented as a spiritual officeholder risks confusing access to language with the responsibilities of ministry. That is this article's Christian application, not a claim that the biblical passage foresaw chatbots or settled every possible use.
What a meaningful next test would establish
The most useful next step is a reviewable trial with a clearly bounded purpose. Before involving learners, researchers could test whether answers cite relevant material, whether disagreements remain visible and whether out-of-scope questions are consistently redirected. Human reviewers would need to assess failures as well as successful demonstrations.
The human side also needs testing. A referral has little practical meaning unless someone can receive it, understand what follow-up is authorized and respond within the institution's actual capacity. The people involved should know those arrangements before they rely on the service.
These are implications of the evidence assembled here, not features verified in a deployed system. The new publication offers a framework for asking better design questions. Evidence that its proposed safeguards work remains a separate task. Keeping that distinction visible is how a promising idea becomes something an educational community can responsibly evaluate.