Different Saju Results? Match the Input Before the Comparison
For AI human review, the first decision is whether one small example can be reviewed by a person. The same birthday on different saju screens is not enough to establish that the inputs match. Before comparing results, record the calendar selection, birth time, location-related settings, and any calculation options the demos expose. Compare the underlying chart separately from the written interpretation. Matching results show agreement on that case; they do not establish calculation accuracy.
Match the conditions: use a fictional birthday and preserve every selected setting.
Separate the outputs: compare chart fields before comparing interpretation text.
Limit the conclusion: record agreement, disagreement, or uncertainty without declaring a calculator correct.
This field-test worksheet is prepared on 2026-09-08. Tested date: not available. The supplied evidence contains no completed demo comparison, captured birth inputs, or resulting charts. The conditions below are a reproducible test plan, and the result cells remain unobserved.
The birthday is only the beginning of the record
“Same birthday” describes what a person remembers entering. A useful comparison needs a record of what each interface actually accepted.
A calendar selector, a time field, and an advanced-settings panel belong in that record. If a demo presents a conversion or adjusted value after submission, preserve that too. The submitted value and the displayed value are separate evidence.
For this exercise, call the interfaces Demo Alder and Demo Birch. These are fictional labels, not reviewed products. Use a fictional birthday rather than personal information.
The aim is narrow: establish whether the available evidence supports a meaningful comparison. It is not a test of whether an interpretation describes someone convincingly.
A shared birthday is a starting point, not a complete comparison record.
Put the input contract on the page
The following table is the reusable artifact. Copy it into an existing text document or spreadsheet; the worksheet requires no purchase. Access conditions for any demo remain a separate question.
Replace bracketed entries before running the comparison. Reuse the same fictional date and exact clock time wherever those fields are supported. Do not treat blank cells as matching values.
| Condition | Fictional case to prepare | Demo Alder receipt | Demo Birch receipt |
|---|---|---|---|
| Entered birth date | [YYYY-MM-DD], explicitly fictional | Not observed | Not observed |
| Calendar selection | [Exact calendar label] | Not observed | Not observed |
| Leap-month selection | [Selected value or not applicable] | Not observed | Not observed |
| Birth time | [HH:MM], with format recorded | Not observed | Not observed |
| Time certainty | Known synthetic time; avoid an unknown-time default | Not observed | Not observed |
| Birthplace, if requested | [Same selected place] | Not observed | Not observed |
| Timezone setting, if exposed | [Exact label or offset] | Not observed | Not observed |
| Clock adjustment, if exposed | [Exact option and selected value] | Not observed | Not observed |
| Date-boundary convention, if documented | [Quoted setting label or documentation reference] | Not observed | Not observed |
| Other required selectors | [Field labels and selected values] | Not observed | Not observed |
| Submitted or converted values shown | Capture separately from entered values | Not observed | Not observed |
| Demo revision and capture date | Record whatever version information is available | Not observed | Not observed |
This is an inspection list, not a claim that every demo implements these options or that every option changes every output.
Use “not exposed” when a control cannot be found. Use “undocumented” when its behavior cannot be established from the available explanation. Those labels preserve uncertainty; neither means “same as the other demo.”
Artifact caption: Input-condition worksheet for a fictional birthday. The receipt columns are unobserved because no demo captures accompany this article.
Capture what survived submission
Begin with a fresh input form. Enter the fictional case, inspect the selected settings, and capture the form before submission. Then preserve the result screen and any summary of the accepted inputs.
This creates a trace from intended values to submitted values to displayed output.
If the result screen omits the input summary, retain the form capture alongside it. Record that the result itself does not confirm the accepted values. Avoid reconstructing settings later from memory.
For a developer evaluating a source package, this trace is useful evidence of what the demo makes inspectable. It does not establish how the underlying implementation handles a setting that the interface hides.
Choose a straightforward baseline before exploring ambiguous or boundary-sensitive cases. That is a recommendation for easier diagnosis, not a claim that any baseline has passed here.
Compare the chart before the prose
Keep the output record separate from the input table. Otherwise, “different result” can become an untidy mixture of a changed selector, a changed chart, and a differently worded paragraph.
| Output layer | What to preserve | What the comparison can establish |
|---|---|---|
| Accepted-input summary | Exact displayed date, time, and calendar information | Whether the displayed inputs agree |
| Core chart | Corresponding year, month, day, and hour fields, where shown | Whether those displayed fields agree |
| Additional calculations | Field labels, values, and available definitions | Whether comparable named outputs agree |
| Written interpretation | Relevant passages linked to the displayed chart | Whether the wording or stated interpretation differs |
| Warnings and omissions | Unknown-time notices, missing fields, validation messages | What limits the comparison |
Preserve original labels before mapping fields across interfaces. Similar placement is not sufficient evidence that fields mean the same thing.
If the charts match while the prose differs, record that exact finding. Do not turn a writing difference into an alleged calculation defect. If the charts differ, identify the affected fields before speculating about the explanation.
Different prose and different chart values are different findings.
Change a condition only after saving the baseline
Once the baseline receipts are complete, duplicate the case and change a single exposed condition. Keep the remaining recorded inputs fixed.
A calendar selector is a possible variable if both interfaces provide it. An unknown-time option is another possible variable. These are proposed checks, not observed failures.
For every variation, record:
- The field changed and its previous and new values.
- The input summary displayed after submission.
- The output fields that changed or remained unchanged.
- Any warning, rejected submission, or missing result.
- The settings whose behavior remains unknown.
Changing a selector while keeping the typed date fixed tests the response to that selector. It does not necessarily preserve the same intended birth event. State which comparison you are attempting.
If several conditions change together, preserve the result but mark the cause unresolved.
Agreement is a smaller claim than accuracy
Suppose a future run produces matching chart values. The supported conclusion would be: these interfaces displayed matching values for this recorded case under the visible conditions.
That would not establish independent correctness. Agreement leaves open whether the implementations use the same assumptions or share a mistake.
A stronger calculation check needs an independently justified expected result, with the relevant conventions stated. No such reference calculation is supplied here. Nor would a correct calculation establish that a written prediction is true.
Human review should therefore inspect the evidence trail, not reward an interpretation for sounding personally convincing. Whether interpretation text is AI-generated or otherwise produced, its fluency is not a calculation receipt.
Agreement is evidence of consistency; accuracy needs a justified reference.
Leave missing evidence visible
There is no verified failed submission, mismatched chart, or successful comparison to report in this article. Inventing one would defeat the worksheet’s purpose.
The present limitation is specific: there are no captured demo inputs and outputs from which to draw a test conclusion. Hidden settings, unclear field definitions, and missing input summaries are potential obstacles to document during a future run, not defects already observed.
The final decision is to use this worksheet as a prerequisite for comparing demos. A matched, documented case can support a bounded consistency finding. An undocumented case should remain unresolved.
Use the worksheet alongside the source-pack page to assess the linked demo and its stated installation conditions.
Related build logs
- Human Review for Marketing Percentages That Leave Out the Comparison
- First AI SOP Template: Put the Failed Example First
Match and capture the input conditions, compare chart fields separately from prose, and treat matching results as agreement rather than proof of accuracy.