GENIE X HUBPROJECT: ECLIPSE
EVALUATION FRAMEWORK

Official Judging System.

PROJECT: ECLIPSE uses a strict 100-point evaluation breakdown focused on autonomous reasoning, tool usage, system reliability, and working execution.

Evaluation Criteria & Weights

TOTAL: 100%
Problem Understanding & Relevance10%

Does the team clearly articulate the problem domain and why their solution addresses it?

Agentic Reasoning & Planning20%

Does the system demonstrate autonomous multi-step reasoning, decomposition and planning — not just prompt-response?

Tool / API / System Integration15%

Does the agent call real tools, APIs or systems and act on the results?

Autonomous Task Completion20%

Does the system complete meaningful end-to-end tasks without constant human prompting?

Reliability & Handling Failure Cases15%

How well does the system handle unexpected inputs, API failures, ambiguous data and edge cases?

Innovation10%

Does the solution introduce a novel approach, architecture or application of AI/agentic systems?

UX / Demonstration10%

Is the demonstration clear, compelling and does it show the system working end-to-end?

DEMONSTRATION EXPECTATIONS

What Judges Expect to See

01

INPUT

System receives a realistic real-world input

02

PLAN

Agent decomposes the task and selects tools

03

TOOL USE

APIs, databases or systems are queried

04

DECISION

Agent reasons over results and decides next action

05

ACTION

Agent executes or proposes a real-world action

06

VERIFICATION

Outcome is checked, failures handled, human alerted if needed

RELIABILITY MATTERS (15%)

AI agents operate in unpredictable environments. Submissions will be explicitly tested against edge cases, API rate limits, malformed inputs, and execution failures. Systems that gracefully handle failures and self-correct will earn top marks.