---
title: I Built One Psych Test, and It Became Twenty — A Structure Where Adding One Test Means Adding One File
url: https://oosioo.com/en/p/%EC%8B%AC%EB%A6%AC%EA%B2%80%EC%82%AC-%ED%95%98%EB%82%98%EB%A5%BC-%EB%A7%8C%EB%93%A4%EC%97%88%EB%8B%A4%EA%B0%80-%EC%8A%A4%EB%AC%B4-%EA%B0%9C%EA%B0%80-%EB%90%90%EB%8B%A4-%EA%B2%80%EC%82%AC-%ED%95%9C-%EC%A2%85%EC%9D%84-%ED%8C%8C%EC%9D%BC-%ED%95%98%EB%82%98%EB%A1%9C-%EB%8A%98%EB%A6%AC%EB%8A%94-%EA%B5%AC%EC%A1%B0
date: 2026-10-02T10:33:42+00:00
author: SYSOP
summary: It started as a casual quiz, just for fun. Then I rebuilt it into a framework where dropping in one file for one test automatically wires up scoring, reports, and screens — and it grew to twenty. The real power of a click isn't knocking out one thing fast; it's setting up the board so the twenty-first comes for free. But looking at how this is done elsewhere, I found that making something that looks like a test and making an actual test are entirely different leagues of difficulty.
---
# I Built One Psych Test, and It Became Twenty — A Structure Where Adding One Test Means Adding One File

I made something like a personality test, just for fun. You answer some questions, it scores them, and spits out a report along the lines of "you're this kind of person." My family tried it and it was surprisingly fun. It's not a real clinical test — just a casual mirror for looking at your own tendencies.

That got me greedy. I wanted to make more tests like it — not just personality, but interests, strengths, relationship tendencies, that sort of thing. The problem is that if you build each test by copying the whole thing and swapping out the questions, your hands give out by around the tenth one. The moment you're managing scoring logic, report generation, and screens separately for every single test, you're in hell.

So I changed direction. I restructured things so that one test could be built as one file.

> It looks like the number of tests is going to keep growing. Instead of copying code for
> every test, refactor this so that defining one test in one file automatically wires up
> the listing, scoring, report, and screen. When I add a new test, I want there to be
> exactly one file I touch.

Right now there are about twenty similar casual tests. They're split into a few categories — personality, interests, relationships, work style, study style, plus a few made just for fun.

![A diagram showing how dropping one test in as a single file automatically wires up the listing, scoring, report, and screen](/uploads/cc3a967fcbfa3bb7.webp)

## One File, One Test

The process for adding a new test shrank down to this:

- Make one file in the folder. In it, write the test's name, the dimensions it measures, the questions, and the report settings.
- Register that file with one line in the list.
- Done.

Scoring is handled by a shared engine that works regardless of how many dimensions there are, report sentences are generated by a shared prompt, and the screen is drawn by a shared template. Whatever differs between tests all lives in that one file, and everything else is shared. When I needed to write hundreds of new questions, I ran several AIs in parallel, had them split up the dimensions, and merged and verified everything at the end.

The key here was splitting the work in two. Code calculates the scores. The rules are fixed, so there's no reason to hand that to AI — if you do, the answer actually changes every time. Instead, turning those scores into a story a human can read is the AI's job. Calculation goes to the machine, narration goes to the language model. Keep that boundary, and the results stay stable.

![A diagram showing the division of labor: code handles score calculation, AI handles the narrative report](/uploads/8746d4fd10a3f8d2.webp)

## An Actual User Changed the Structure Again

I gave it to an acquaintance to try, and they said, "All the questions are phrased as obviously good things, so I just end up marking 'agree' on everything."

They had a point. Not many people would say "no" to something like "I have a strong sense of responsibility." Since everyone wants to present themselves favorably, the scores skew upward.

So I tore it down and rebuilt it again. I changed most of the tests to a forced-choice format. When you put two different tendencies side by side and ask which one is closer to you, you can't rate both favorably at once. That said, I kept the original scale format for a handful of tests — things like fatigue level or resilience, where the absolute degree of severity actually matters. And because each test was already self-contained in a single file, changing the format meant touching only that one file.

![A screen showing the switch from scale-based to forced-choice format to prevent the pull toward looking good](/uploads/a657e41505f75d78.webp)

## Turns Out, There's Already a Lot of This Out There

Once I'd built it, I got curious and looked around. It turns out there are a lot of places doing something similar.

The most famous branch of this is the kind that sorts personality into a four-letter type. Just saying "four-letter type" is probably enough for most people to guess what I mean. The most widely used online version is run by a British company, and by their own count it's been taken over a billion times. The interesting part is that while it borrows the format of that famous four-letter test, underneath it actually leans on a different model that's better regarded in academia — the one commonly called the "Big Five."

The original four-letter test itself gets a complicated reception. It was created in the early 20th century by a mother-daughter pair basing their work on Jung's theory, and neither of them had formal training in psychology. Academia has leveled sharp criticism at it for a long time. It's common for the same person to get a different type if they retake it a few weeks later, and back in 1991 the U.S. National Academy of Sciences stated flatly that there wasn't sufficient grounds to use it for career counseling. The Big Five model that researchers trust more has its questions published openly — a pool of thousands of public items that anyone can pick up and use for research.
These days, using AI to evaluate people has also grown quite a bit. One company used AI to analyze applicants' video interviews for hiring. For a while it even read facial expressions, but they eventually dropped that facial analysis — the reasoning being that it had no real value and drew nothing but controversy. There's also a service that claims to infer someone's personality just from public social media data. But when outside researchers took it apart in 2022, they found the personality score changed depending on whether the exact same resume was submitted as a PDF or as plain text. The researchers concluded it "cannot be considered a valid testing instrument." As cases like these piled up, New York City and the EU began requiring mandatory bias audits specifically for AI used in hiring.

On the opposite end, there's an entirely different world: places that build real tests. Even just within Korea, there are organizations that have imported and standardized the Korean versions of famous overseas tests, and others that publish over 260 kinds of psychological assessments. Producing even one such test involves establishing norms against thousands of subjects, and statistically verifying both that retaking it yields similar results and that it actually measures what it's meant to measure. Question sets running into the hundreds are written with dozens or hundreds of professors and clinicians involved. That's a completely different starting line from questions I cooked up with AI over a weekend.

At the same time, as that four-letter test became hugely popular in Korea, it also started getting used in places it was never meant for. There have even been job postings explicitly telling certain types not to apply. Experts are unanimous in warning against this — using it to judge personality in hiring is dangerous. The moment you load someone's pass-or-fail outcome onto a tool that was built just for fun, it becomes an entirely different problem.

## So, Why Does This Count as a Click

People tend to think of a "click" as knocking out one thing in an instant. But the real source of power lies elsewhere: setting up the board so the twenty-first thing comes for free. AI happens to be especially good at making one more of the same shape of thing. So if, from the start, you build not a single test but a frame to hold tests, then afterward, every time a new test idea comes to mind, all you have to do is drop in one file. With the effort of building one, you end up holding a board with twenty.

But after looking around, one more thing became clear. Making something that looks like a test and making an actual test are completely different levels of difficulty.

![A comparison of the difficulty gap between building the appearance of a test and becoming a genuinely standardized test](/uploads/4b597d3d333f2965.webp)

Once you've built the frame, you can stamp out the surface as much as you like. Questions appear, scores come out, a plausible-looking report gets attached. A click gets you this far quickly. But whether that score is actually measuring something, whether it comes out the same way the next time you take it, and whether it's fit to be used in judging a person — that's an entirely different domain. That's not code; it's decades of accumulated knowledge from the field of psychometrics. Norms drawn from thousands of people, reliability and validity verification, the judgment needed to write and select items. That's not something that comes out just because you tell AI, "make me a test."

That's why, from the start, I pinned down what I'd made as something made for fun. It's a light mirror for looking at yourself, not a diagnosis. What a click opens up is broad — on your own, you can quickly stand up something fairly convincing. But the moment it carries the weight of a real person on it — hiring, diagnosis, selection, that sort of thing — it absolutely needs the hands of someone who actually knows the field. Anyone can now hold the tool, but knowing what it's actually measuring is still the expert's job.

To add one honest note: maintenance after building it is also its own separate matter. When users pile up, the AI calls that generate reports can back up in a queue, and I had to bolt on a separate mechanism just to process them in order. Even something born from a click still needs tending to as long as it's alive.

