johba/harb - Forgejo: Beyond coding. We forge.

johba/harb

Author	SHA1	Message	Date
openhands	c42a1ca768	fix: evo_run004_champion fitness inflated by token value (#670 ) (#704 ) - Add fitness_flags="token_value_inflation" to evo_run004_champion in manifest.jsonl so callers can detect the inflated value without discarding the entry entirely. - Add effective_fitness() helper in evolve.sh pool admission (step 5) that returns 0 for any entry with a token_value_inflation flag, preventing inflated scores from biasing the top-100 evolved pool ranking or eviction decisions. - Document in evolve.sh that raw fitness values are only comparable within the same evaluation run. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-14 01:08:13 +00:00
openhands	b7d0b63ca1	fix: int(e.get('fitness', 0)) crashes on null-fitness manifest entries (#711 )	2026-03-13 23:37:00 +00:00
johba	0c4cd23dfa	fix: feat: Seed kindergarten — persistent top-100 candidate pool (#667 ) (#683 ) Fixes #667 ## Changes ## Summary Implemented persistent top-100 candidate pool in `tools/push3-evolution/evolve.sh`: ### Changes `--run-id <N>` flag (line 96) - Optional integer; auto-increments from highest `run` field in `manifest.jsonl` when omitted - Zero-padded to 3 digits (`001`, `002`, …) Seeds pool constants (after path canonicalization) - `SEEDS_DIR` → `$SCRIPT_DIR/seeds/` - `POOL_MANIFEST` → `seeds/manifest.jsonl` - `ADMISSION_THRESHOLD` → `6000000000000000000000` (6e21 wei) `--diverse-seeds` mode now has two paths: 1. Pool mode (pool non-empty): random-shuffles the pool and takes up to `POPULATION` candidates — real evolved diversity, not parametric clones 2. Fallback (pool empty): original `seed-gen-cli` parametric variant behavior - Both paths fall back to mutating `--seed` to fill any shortfall Step 5 — End-of-run admission (after the diff step): 1. Scans all `generation_*.jsonl` in `OUTPUT_DIR` for candidates with `fitness ≥ 6e21` 2. Maps `candidate_id` (e.g. `gen2_c005`) back to `.push3` files in `WORK_DIR` (still exists since cleanup fires on EXIT) 3. Deduplicates by SHA-256 content hash against existing pool 4. Names new files `run{RUN_ID}_gen{N}_c{MMM}.push3` 5. Merges with existing pool, sorts by fitness descending, keeps top 100 6. Copies admitted files to `seeds/`, removes evicted evolved files (never hand-written), rewrites `manifest.jsonl` Co-authored-by: openhands <openhands@all-hands.dev> Reviewed-on: https://codeberg.org/johba/harb/pulls/683 Reviewed-by: review_bot <review_bot@noreply.codeberg.org>	2026-03-13 20:45:03 +01:00
johba	3f435f8459	fix: evolution scoring — 3 bugs made all candidates report fitness=0 (#665 ) ## Three bugs in evolve.sh 1. Heredoc stdin conflict — `py_stats()` used `<<PYEOF` heredoc which stole stdin from the pipe, so python never received score values → stats always `min=0 max=0 mean=0` 2. Bash integer overflow — global best comparison used `[ $MAX -gt $GLOBAL_BEST_FITNESS ]` which overflows on uint256 wei values (>9.2e18) → best always tracked as 0 3. candidate_id mismatch — evolve.sh looked up `gen0_c000` but batch-eval produces `candidate_000` (derived from filename) → score lookup always returned default 0 All 3 previous evolution runs (150+ candidates) reported all zeros despite batch-eval correctly scoring them at ~8.26e21 wei. ## Fix - `py_stats`: heredoc → `python3 -c` inline - Global best: bash `[ -gt ]` → `python3` big number comparison - Score lookup: use `basename $CAND_FILE` instead of synthetic CID Co-authored-by: root <root@debian-g-2vcpu-8gb-ams3-01> Reviewed-on: https://codeberg.org/johba/harb/pulls/665 Reviewed-by: review_bot <review_bot@noreply.codeberg.org>	2026-03-13 10:02:24 +01:00
openhands	89a2734bff	fix: address review findings for diverse seed population (#638 ) - evolve.sh: fix fail-in-subshell bug — run seed-gen-cli as a direct command so its exit code is checked by the parent shell and fail() aborts the script correctly; redirect stderr to log file instead of discarding it with 2>/dev/null - seed-generator.ts: reorder enumerateVariants() to put STAKED_THRESHOLDS outermost (192 entries/block) so that selectVariants(6) with stride=192 covers all 6 staked% thresholds; remove false doc claim about "first variant is current seed config"; add comments explaining CI=0n is intentional in all presets - seed-gen-cli.ts: emit a stderr diagnostic when count exceeds the 1152-variant cap so the cap is visible rather than silently producing fewer files than requested - test: strengthen n=6 test to assert all STAKED_THRESHOLDS values are represented in the selected variants Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-13 05:21:05 +00:00
openhands	850131b74f	fix: feat: Push3 evolution — diverse seed population (#638 ) Add seed-generator.ts module and seed-gen-cli.ts CLI that produce parametric Push3 variants for initial population seeding. Variants systematically cover: - Staked% thresholds: 80, 85, 88, 91, 94, 97 - Penalty thresholds: 30, 50, 70, 100 - Bull params: 4 presets (aggressive → mild) - Bear params: 4 presets (standard → very mild) - Tax distributions: exponential (seed), linear, sqrt Total combination space: 6×4×4×4×3 = 1152 variants. selectVariants(n) samples evenly so every axis is represented. evolve.sh gains --diverse-seeds flag: when set, gen_0 is seeded with parametric variants instead of N copies of the same mutated seed. Remaining slots (if population > generated variants) fall back to mutations of the base seed. All generated programs pass transpiler stack validation (33 new tests). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-13 04:48:04 +00:00
openhands	64f1af3041	fix: feat: Push3 evolution — elitism (top N survive unchanged) (#640 ) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-12 22:29:23 +00:00
openhands	26b8876691	fix: feat: revm-based fitness evaluator for evolution at scale (#604 ) Replace per-candidate Anvil+forge-script pipeline with in-process EVM execution using Foundry's native revm backend, achieving 10-100× speedup for evolutionary search at scale. New files: - onchain/test/FitnessEvaluator.t.sol — Forge test that forks Base once, deploys the full KRAIKEN stack, then for each candidate uses vm.etch to inject the compiled optimizer bytecode, UUPS-upgrades the proxy, runs all attack sequences with in-memory vm.snapshot/revertTo (no RPC overhead), and emits one {"candidate_id","fitness"} JSON line per candidate. Skips gracefully when BASE_RPC_URL is unset (CI-safe). - tools/push3-evolution/revm-evaluator/batch-eval.sh — Wrapper that transpiles+compiles each candidate sequentially, writes a two-file manifest (ids.txt + bytecodes.txt), then invokes FitnessEvaluator.t.sol in a single forge test run and parses the score JSON from stdout. Modified: - tools/push3-evolution/evolve.sh — Adds EVAL_MODE env var (anvil\|revm). When EVAL_MODE=revm, batch-scores every candidate in a generation with one batch-eval.sh call instead of N sequential fitness.sh processes; scores are looked up from the JSONL output in the per-candidate loop. Default remains EVAL_MODE=anvil for backward compatibility. Key design decisions: - Per-candidate Solidity compilation is unavoidable (each Push3 candidate produces different Solidity); the speedup is in the evaluation phase. - vm.snapshot/revertTo in forge test are O(1) memory operations (true revm), not RPC calls — this is the core speedup vs Anvil. - recenterAccess is set in bootstrap so TWAP stability checks are bypassed during attack sequences (mirrors the existing fitness.sh bootstrap). - Test skips cleanly when BASE_RPC_URL is absent, keeping CI green. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-12 11:54:41 +00:00
openhands	ade7e2033a	fix: Evolution pipeline UUPS upgrade + Foundry PATH (#593 ) - Add virtual to Optimizer.calculateParams() for UUPS override - Create OptimizerV3.sol: UUPS-upgradeable optimizer with transpiled Push3 logic - Update deploy-optimizer.sh to deploy OptimizerV3 instead of Optimizer - Add ~/.foundry/bin to PATH in evolve.sh, fitness.sh, deploy-optimizer.sh	2026-03-12 06:47:35 +00:00
openhands	0496c94681	fix: address review findings in evolve.sh (#546 ) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-11 22:06:18 +00:00
openhands	2ee7feb621	fix: address review findings in evolve.sh (#546 ) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-11 21:29:14 +00:00
openhands	547e8beae8	fix: Push3 evolution: selection loop orchestrator (#546 ) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-11 20:56:19 +00:00

12 commits