Researcher guide

Developing Your Own Framework

DeCiphr is designed to strictly apply the analytical framework you give it. This is a step-by-step guide for developing a coding scheme the AI can apply consistently.

What a framework is in DeCiphr

A framework is a self-contained coding scheme: a set of codes with precise definitions, an optional controlled vocabulary of values for each code, and a master prompt that situates the analysis within your analytical tradition. It is stored as a .sigframe file — plain, human-readable JSON — so it can be reviewed, edited, shared with collaborators, and reported in your research output.

DeCiphr runs in strict coding mode only. The AI assigns values from the codes you define; it does not introduce categories, terminology, or observations beyond your framework. Where the visual evidence does not warrant a confident assignment, it returns Uncodable rather than guessing. You review and can edit every result.

Every framework can be organized into two layers that do different work. The AI layer holds deductive codes assignable from what is visually present in the image: languages, layout, materiality, composition. The researcher layer holds interpretive questions that require fieldwork presence, community knowledge, and ethnographic judgment. The AI never answers researcher-layer questions, by design.

The elements of a framework

Identity

Name, a short description of what the framework analyzes, and the field or discipline it belongs to.

Master prompt

The framing instruction the AI receives before your codes: the analytical role it takes, how to treat text and image, and the required output discipline (values only, Uncodable when evidence is insufficient).

Codes

The heart of the framework. Each code has a name, a definition written as a coding instruction, and optional sub-categories that act as a controlled vocabulary of permissible values.

Layer assignment

Each code belongs to the AI layer Researcher layer or both. Layer assignment is set per code in the web app editor; codes created without a layer setting are treated as AI-layer codes.

References & attribution

The published sources your framework operationalizes, in your citation style, plus creator details (name, institution) so shared frameworks remain attributable.

Visibility

Frameworks are private by default. You can share them as .sigframe files directly, or publish them to the in-app Framework Library for other researchers to browse and import.

Building a framework, step by step

1

Open the framework editor

In the web app (app.deciphr.tech) or the mobile app, go to Frameworks and create a new framework. The web app editor includes the per-code layer picker, so it is the better place to build two-layer frameworks.

2

Define the identity

Give the framework a name, a one- or two-sentence description of what it analyzes, and its field or discipline. Name deliberately — include a version (for example, "v1") so revisions stay distinguishable.

3

Write the master prompt

State the analytical role, the tradition the codes come from, whether the AI should attend to text, image, or both, and the output rules: assign only defined values, no prose, Uncodable when the image provides insufficient information.

4

Add your codes

One code per analytical question. Each needs a name, a definition written as an instruction, and — wherever the theory provides them — sub-categories as the controlled set of permissible values. On mobile you can also import codes from a CSV file.

5

Assign each code a layer

In the web app editor, mark each code AI, Researcher, or Both. Ask one question of every code: can this be answered from the image surface alone? If it requires being there, knowing the community, or interpreting significance, it belongs to the researcher layer.

6

Pilot, refine, then share

Run the framework on a small batch of five to ten items before committing to a corpus. Where the AI's assignments diverge from your intent, refine your codes and definitions. When stable, export the .sigframe or publish to the Framework Library.

Methodological and Practical Considerations

The definition is the instruction. The AI sees only what you write; it is instructed by your framework at the moment of analysis, not trained on it, and it has no access to the source text behind your codes. Definitions therefore need to be self-contained: state what to look for, in what order, and what form the answer takes. Write in the imperative ("Identify every language present…", "Assign one of the following…") and give decision rules for hard cases rather than leaving them to judgment.

Constrain values wherever possible. Sub-categories turn a code from an open description into a controlled assignment, which is what makes coding consistent across hundreds of items and comparable across coders. Reserve open-ended codes for the researcher layer, where your own interpretive writing belongs.

Draw the layer boundary clearly. The AI layer is for structural, surface-readable features; it returns coded values only, not interpretive prose. Anything that depends on fieldwork presence, community knowledge, policy context, or reflexive judgment belongs to the researcher layer. The division is what preserves your interpretive authority.

Plan for Uncodable. A well-built framework tells the AI what to do when evidence is insufficient. A high Uncodable rate in your pilot is diagnostic: either the definition demands information the images cannot supply, or the code belongs in the researcher layer.

One code, two layers

The example below is taken from the Geosemiotics framework (Scollon & Wong Scollon, 2003) used in the developer's own fieldwork. It shows how a single construct, code preference, is split into an AI-layer structural code and a set of researcher-layer interpretive questions.

Master prompt
"You are a geosemiotics structural coding assistant. Analyze this sign image using the 7 codes below… For each code, assign the controlled value that matches what is directly visible in the image — consider both text and image elements. Do not write descriptions, evidence, or sentences. Do not interpret community meaning or require fieldwork knowledge… For any code where the image provides insufficient information, return the string "[UNCODABLE]" as the value."

Note what the prompt does: it names the analytical role, fixes the scope to what is visible, forbids interpretation, and defines the insufficient-evidence behavior.

AI layer

Code Preference

"Identify every language, script, or visual code present on the sign — in both text and image. Determine which is the preferred (dominant) code by identifying the spatial axis organizing privilege: Center–Margin…; Top–Bottom…; Left–Right…; Earlier–Later…. State the axis, name the dominant and subordinate codes explicitly, and describe the positional or visual evidence… If only one code is present, state it. [Scollon & Scollon 2003, Ch. 6, pp. 116–125]"
Sub-categories (controlled values)
Center–MarginTop–BottomLeft–RightEarlier–Later

Everything here is answerable from the image surface: which codes appear and how they are spatially arranged. The definition carries the framework's own terms and its citation, so the coded output remains traceable to the source.

Researcher layer

The paired interpretive questions

  • Does the dominant code index the local community or symbolize something non-local? Values: Indexes Local · Symbolizes Non-Local · Mixed · Unclear
  • Is this code arrangement shaped by law, policy, or institutional mandate at this site? Values: Yes: Government · Yes: Institutional · No · Unknown
  • Why is this code preferred here? What social, political, or economic forces produced this arrangement? Open-ended — answered from fieldwork.

None of these can be read off the image. Whether a dominant language on a shopfront indexes the surrounding community or symbolizes something else entirely is knowable only through presence at the site — which is precisely why these questions are reserved for the researcher.

A short checklist

Every code answers exactly one analytical question, in the framework's own terms.
Every definition is self-contained and imperative; it should make sense to a coder who has never read the source text.
Codes with theory-given categories have sub-categories; open description is reserved for the researcher layer.
Every code is assigned a layer, and nothing in the AI layer requires fieldwork knowledge to answer.
The master prompt states the role, the scope (text, image, or both), and the Uncodable rule.
Piloted on five to ten items, definitions revised where assignments diverged, before coding the corpus.

Questions about framework design, or a framework you would like to share? Get in touch.