The practical answer

Score the work a product can demonstrate with your test cases: employer separation, one authoritative transmittal, reconciled counts, approval history and usable filing evidence. Keep essential requirements separate from the weighted score so a polished demo cannot compensate for a missing control.

This scorecard is for employers comparing the transmittal controls in 1094-C software. It is an original evaluation method, not an IRS certification or a ranking of named vendors. Use fictional data, the same tasks and the same rating scale for every candidate. The form references below use the final 2025 instructions.

Define a passing result before scheduling the demonstration

Begin with three or four situations your reporting team actually handles: multiple batches for one employer, two employers in a group, revised monthly counts, or a correction after acceptance. Write the expected result and the artifact you must be able to retrieve. Give vendors the task in advance, then ask your own operator to repeat it during the session.

For example, the 2025 instructions for line 19 require one authoritative transmittal for each ALE Member. An evaluation task can therefore ask how the product identifies that transmittal when an employer has several batches. The instruction establishes the reporting requirement; it does not require a particular screen or automated warning.

State whether an acceptable result may involve a documented manual review. A product limitation and a workable operating procedure are different findings, and both should be visible in the scorecard.

Use weights tied to transmittal work

Adjust the following weights before seeing a product demonstration. Their purpose is to express your priorities consistently. They total 100, and the evidence requests make the categories harder to satisfy with a feature name alone.

Example 1094-C software evaluation weights
CapabilityWeightDemonstration evidence
Entity and authoritative control25Two employers, several batches, and a clear authoritative selection for each employer.
Count reconciliation25Trace a batch count, employer total and monthly count to their separate sources.
Evidence export20Retrieve the approved version and identifiable submission evidence outside the demo screen.
Correction history15Locate a prior accepted transmittal and show how a later correction remains associated.
Approval controls15Show who approved which version and what happens after a material edit.

These are transmittal-specific priorities. Contract terms, price, general security review and employee coding support can have separate evaluations. Combining everything into this table makes it difficult to see whether the product solves the 1094-C problems you brought to the demonstration.

Rate observed evidence, with limitations beside the number

Use a five-level scale: 0 means not demonstrated; 1 means a verbal or written capability claim; 2 means the vendor demonstrated the task; 3 means your operator repeated it; 4 means your operator repeated it and retained the requested evidence. A score of zero can mean the session ran out of time, so record the reason rather than treating every zero as a proven product defect.

For each category, calculate weight multiplied by rating divided by 4. Attach the product version, demonstration date, dataset version, observed result and any manual steps. If a category contains several tasks, establish its scoring rule beforehand. For example, require all listed tasks to reach a level before awarding that rating.

Do not award evidence-export credit merely because a vendor opens a receipt on screen. Ask the evaluator to retrieve and interpret the artifact. Publication 5165 distinguishes receipt evidence from processing acknowledgment, which makes that distinction worth testing.

Worked example: a score of 72.5 with a remaining requirement

Fictional example: Westward Supply evaluates an unnamed product using the weights above. The team awards ratings of 4 for entity controls, 3 for counts, 2 for evidence exports, 3 for correction history and 2 for approvals.

The weighted points are 25, 18.75, 10, 11.25 and 7.5. Their sum is 72.5 out of 100. The export task earned only a 2 because the vendor demonstrated it but the employer's operator could not repeat it during the session. That is a specific unresolved observation, not proof that export is impossible.

Westward had separately defined operator-accessible historical acknowledgments as essential. It therefore records the evaluation as incomplete despite the 72.5 score and requests a focused retest. A higher aggregate score would not remove that requirement. The team keeps the first result and adds dated retest evidence so the decision remains explainable.

Keep essential requirements outside the average

Write a short list of requirements that cannot be traded against other features. Examples might include support for a specific prior reporting year, separation of two employer entities, a retrievable authoritative transmittal, or a viable export when the relationship ends. Identify who set each requirement and what evidence will satisfy it.

Phrase these as your organization's acceptance conditions, not universal legal requirements for software. If the product relies on vendor assistance, specify the request process and expected response in the evaluation record. A demonstrated service-assisted workflow may be acceptable, but its dependency should not disappear into a checkbox labelled supported.

Ask about the migration of historical reporting evidence before declaring the evaluation complete. A product may prepare a new year successfully while leaving prior accepted returns in another system.

Turn gaps into reproducible follow-up tasks

Each unresolved item should become a small retest: dataset, starting state, operator action, expected result and evidence to retain. For a count discrepancy, request the exact source-to-form reconciliation. For an approval question, edit a previously approved transmittal and observe whether the approval remains valid, is withdrawn, or requires a manual process.

Keep subjective usability notes beside the score. Record which steps confused the operator and how long a repeated task took, without inventing a universal speed threshold. A narrow scorecard can support a considered selection, but it cannot establish that the employer's source facts or final tax reporting will be correct.

Finish by documenting the selected operating procedure, unresolved dependencies and the person responsible for each. Preserve the actual demo outputs so implementation can be checked against what was shown.

From reporting requirement to demonstrated capability

From reporting requirement to demonstrated capability: Define expected result; Repeat the task; Retain the evidence; Score and resolve gaps
An original purchasing evaluation method. Scores describe observed evidence, not IRS approval or guaranteed reporting accuracy.
Read the workflow as text
  1. Define expected result. Choose employer, count and historical evidence tasks before the demo.
  2. Repeat the task. Have your operator use the same fictional dataset for each product.
  3. Retain the evidence. Save identifiable outputs and record manual dependencies.
  4. Score and resolve gaps. Calculate weighted points while keeping essential requirements separate.

Put this guide to work

1094-C software weighted evaluation worksheet

Save the editable text worksheet and use it with your own records. Keep completed copies in your secure working files.

Download the worksheet TXT

Common questions

Is this an IRS-approved software scorecard?

No. It is an original method for comparing demonstrated transmittal controls. The IRS sources explain reporting fields and AIR procedures. The weights, evidence levels and purchasing gates are practical evaluation choices that an employer should adapt to its own work.

What score should a product receive when a task was not tested?

Record zero for not demonstrated and clearly state that the task remains untested. Do not describe the capability as absent unless the evidence establishes that. Schedule a retest when the task matters to the decision.

Can a manual review satisfy a requirement?

It can if your team explicitly accepts that operating procedure and can perform it reliably. Record the review steps, evidence, responsible role and dependencies. Do not score a manual procedure as an automated prevention feature.

Why separate count reconciliation from entity controls?

A product can keep employers separate while providing little explanation for the numbers on each transmittal. Conversely, good count reports do not prove that the authoritative transmittal is selected correctly. Separate tasks expose these distinct capabilities.

Should we compare only the total score?

No. Compare essential requirements, individual category evidence, unresolved tasks and operating dependencies alongside the total. A small difference in weighted points may be less consequential than a missing historical export or an approval process your team cannot operate.

Official sources and scope

Sources checked September 5, 2026. Use the edition for the tax year and filing method you are working with; later instructions may change thresholds, fields, or procedures.

  1. IRS 2025 Instructions for Forms 1094-C and 1095-C

    Tax year 2025 authoritative transmittal and count requirements used to define demonstration tasks.

  2. IRS Publication 5165, revised December 2025

    AIR receipt, acknowledgment and correction context relevant to evidence demonstrations.