Skip to content

Assessment design

Building fair assessments

A coding test measures two things at once: what the student understands, and how much preparation they could afford. Here is how to shrink the second.

By The Assessly team5 min read

Every technical assessment measures two things at once: what the student understands, and how much preparation they had access to. The second contaminates the first. A student with months of pattern drilling, paid prep material and free evenings will outscore an equally able student who found programming late or studies around a part-time job. Same test, same ability, different score.

Neither student did anything wrong. The test did. And because the paper still looks rigorous (hard problems, strict conditions, a healthy spread of marks), nobody sees the bias in the result. It shows up a year later, when a student who tested well cannot reason about unfamiliar work, and a student who tested badly turns out to be the one the team relies on.

This post is about the choices a faculty member makes while setting a paper, and which of them push the score toward understanding.

1. Choose problems that do not turn on one trick

Most classic algorithm puzzles collapse once you recognise a known technique. Recognising techniques is exactly what high-volume drilling trains, so a puzzle-heavy paper ranks students by how many problems they have already seen. That number tracks time and money.

The fix is to write problems where the obvious approach works and the marks come from doing it carefully: reading the input exactly, handling the empty case, keeping within the limits the constraints imply. A problem like "merge the attendance records from two sections and report students missing from both" rewards clear thinking. A problem that is only solvable by someone who has met the trick before rewards a subscription.

  • Test what the course taught, not trivia or obscure language behaviour
  • Make edge cases the ones real software fails on (empty input, duplicates, boundaries), not adversarial gotchas
  • If a problem has a well-known name, assume the solution is on every phone in the room and change it

2. Give partial credit, and make it honest

An all-or-nothing coding question turns a single off-by-one error into a zero, and the student who understood 90% of the problem scores the same as the one who left it blank. On Assessly the code part of a mark is the share of hidden test cases the submission passes, so a solution that handles the main cases and misses one boundary earns most of the marks and loses the part it got wrong.

That only works if the tests are designed with it in mind. Write tests that each check one thing: a basic case, the empty input, a large input for time limits, the duplicate case. A student then loses marks for the specific thing they missed, and the result tells them what it was.

3. Ask for the reasoning, not just the answer

A faculty member setting a paper on Assessly can require the student to explain their solution in their own words, and decide how much of the mark the explanation carries. The weights are theirs to set; the default is 70% for the code and 30% for the explanation.

This shifts what the paper rewards. A student who writes a straightforward solution and explains it precisely (why the approach works, where it breaks, what they traded away) shows real understanding. A memorised optimal answer cannot explain itself. The write-up behind it tends to be thin, generic, or wrong about its own code.

The explanation should be marked on substance, not vocabulary. "I check every item one by one because the list is not sorted" shows the same understanding as "I perform a linear scan since the input is unordered." One sounds like a textbook. The score should not be able to tell them apart, because polished jargon is one more thing expensive preparation buys.

Memorisation can produce the right code. It cannot produce a precise account of why the code is right.

4. Think twice about negative marking

Negative marking on the aptitude section is meant to stop blind guessing. It also measures something you did not intend to test: appetite for risk. Students who have sat dozens of mock tests learn exactly when a guess pays; students sitting their first proctored paper leave questions blank that they would have got right. If you use it, keep the penalty small, tell students the rule before they start (Assessly shows it on the paper's overview screen), and look at how many questions were left unanswered when you read the results.

5. Let the language be the student's

A student who thinks in Python should not lose marks for being made to write Java. On Assessly the student picks from the languages the problem supports (up to five: Python, JavaScript, Java, C++ and C) and can switch while they work, and the same hidden tests mark every language. Unless the course is specifically about one language, leave the choice with them.

6. Put practice inside the course

Preparation access has a plainer dimension: repetitions. Students from better-resourced backgrounds get more low-stakes attempts before anything counts. Practice problems through the term, marked by hidden tests and with feedback but no grade, move some of those repetitions inside the course instead of leaving them to whoever can pay for them elsewhere. It does not equalise everything. It closes the gap that is cheapest to close.

After the paper: check the paper, not only the students

Fairness is also something you can measure afterwards. When the results come in, look for three things:

  • A question almost everyone failed. That is usually a problem with the question: an ambiguous statement, a test case that expects something the statement never asked for, a time limit too tight for the language most students chose
  • A large gap between sections or branches on one question and not the others. That points at what was taught, not at the students
  • A strong code score with a weak explanation, repeated across many students on one problem. That problem is probably too well known

Tell students in advance that they will be asked to explain their work. That expectation alone changes how they prepare, and it is the cheapest fairness measure there is.

None of this lowers the bar. It aims the difficulty at the thing the grade is meant to certify: whether the student can think through a problem and account for the solution. That is the one kind of difficulty preparation money cannot shortcut.

To see a real submission marked, code and explanation side by side, book a demo.

Written by The Assessly team.

Book a demoHelp centre

Was this page helpful?

Try it on your own batch.

Written by the team that builds Assessly. A 30-minute call shows it working on a real test.