Official repository for:
Co-Evolving Structured Knowledge and Reasoning in Language Models (COLM 2026)
π arXiv β’ π Website β’ π€ Models & Data β’ π» GitHub
Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved information. Structured knowledge bases offer a more controllable alternative, yet they are expensive to construct and often brittle to reason over.
KBevo addresses these limitations with a co-evolving framework that jointly learns to construct a structured knowledge base and reason over it for knowledge-intensive question answering. By optimizing both components end-to-end with QA outcome rewards, reasoning success directly improves the quality of the constructed knowledge base. This leads to larger, better-connected knowledge structures with higher answer reachability, while also improving compositional factual reasoning and controllability compared to standard retrieval baselines.
At inference the same model runs a two-phase policy:
- Phase 1: KB construction. The model reads a collection of input documents and emits
a structured KB of
(entity, relation, value)triples using four special tokens<|db_entity|>,<|db_relationship|>,<|db_return|>,<|db_end|>. The resulting KB is fixed and reused across downstream questions: the paper's evaluation aggregates extracted triples into one retrieval datastore per benchmark. - Phase 2: QA. For each question the model issues structured lookups against the
fixed KB (again via the
<|db_*|>tokens); vLLM stops on<|db_return|>so a nearest-neighbor retriever can splice the KB value back into the running trace.
The examples/single_example.py demo builds a KB from a small
per-example passage set for convenience; that is not the paper eval protocol.
Training is a two-stage pipeline: SFT on 6k Gemini-generated two-phase trajectories,
then GRPO with a factored rollout structure (K=4 phase-1 KBs Γ N=32 phase-2 QAs per
question, M = N/K = 8 QAs per KB) and an F1 outcome reward.
- Released artifacts
- Installation
- Quick Start
- Repository structure
- SFT training
- GRPO training
- Evaluation
- Configuration
- GRPO variants (ablations)
- Citation
Everything below lives in the
KBevo collection
on the Hugging Face Hub. All released model weights are BF16. See
configs/artifacts.yaml for the authoritative manifest.
| Model | Stage | Base | Init | π€ |
|---|---|---|---|---|
KBevo-Qwen3-1.7B-SFT |
SFT | Qwen/Qwen3-1.7B | n/a | π€ |
KBevo-Qwen3-1.7B-GRPO |
GRPO | Qwen/Qwen3-1.7B | 1.7B-SFT (step 368) | π€ |
KBevo-Qwen3-4B-SFT |
SFT | Qwen/Qwen3-4B | n/a | π€ |
KBevo-Qwen3-4B-GRPO |
GRPO | Qwen/Qwen3-4B | 4B-SFT (step 735) | π€ |
SFT dataset. Released as
kilian-group/KBevo-SFT-hotpotqa-6k
(~11.7k full two-phase trajectories over ~5.7k unique HotpotQA train questions,
license Apache-2.0 per the dataset card).
# 1. Clone the repo.
git clone https://github.com/kilian-group/KBevo.git
cd KBevo
# 2. Create the conda env (Python 3.11 + torch 2.8 + vLLM 0.10.2, etc.).
conda env create -f environment.yml
conda activate kbevo
# 3. (Optional) editable install of this repo so the modules under `src/`
# (`agent/`, `data/`, `eval/`, `multi_lmlm/`, ...) are importable without
# manually setting `PYTHONPATH=src`.
pip install -e .
# 4. Copy the site-config template and edit the values in place.
cp configs/cluster.env.example configs/cluster.env
$EDITOR configs/cluster.envconfigs/cluster.env is git-ignored and stores site-specific paths, optional W&B settings,
and optional SLURM resources. Copy the provided template and edit it for your environment;
leaving the W&B variables unset disables W&B logging.
On a SLURM cluster, submit jobs through scripts/submit_slurm.sh, which reads the SLURM_*
settings from configs/cluster.env. Without SLURM, run the canonical scripts directly with
bash.
No hard requirement on FlashAttention. The environment installs xformers==0.0.32.post1
which matches torch==2.8.0; both are Blackwell-ready. Install flash-attn separately only
if you want faster attention on Ampere/Hopper.
The paper's training runs used 4Γ B200 for GRPO and 1Γ B200 for SFT. scripts/train_grpo.sh
ships --gpu_type {B200, H100} presets that adjust hardware knobs only
(per_device_batch_size, vllm_gpu_memory_utilization, accelerate config). Scientific
hyperparameters (learning_rate=5e-6, num_generations=32, etc.) come from the paper
YAML and are constant across GPU types. On other GPUs pass explicit
--per_device_batch_size / --vllm_gpu_memory_utilization overrides.
Scripts run directly; Slurm is optional.
Fastest smoke: 5 examples per dataset on all four benchmarks (~2 min after model download), using a released checkpoint:
bash scripts/eval_kbevo.sh \
--model_path kilian-group/KBevo-Qwen3-1.7B-GRPO \
--datasets hotpotqa,musique,2wiki,popqa \
--num_samples 5Subset evaluation: 1,000 examples per dataset (much slower; useful for a representative signal without the full validation split):
bash scripts/eval_kbevo.sh \
--model_path kilian-group/KBevo-Qwen3-1.7B-GRPO \
--datasets hotpotqa,musique,2wiki,popqa \
--num_samples 1000Paper's full evaluation: one command per dataset at its full validation-split size:
| Dataset | Paper eval size |
|---|---|
| HotpotQA | 7,405 |
| MuSiQue | 2,417 |
| 2WikiMultiHopQA | 12,576 |
| PopQA | 1,399 |
--model_path accepts either a local checkpoint directory or a Hugging Face repo id
of the form owner/name (it materializes the latter to a local snapshot).
Single-example inference demo (no dataset loader required). Defaults to
kilian-group/KBevo-Qwen3-1.7B-GRPO and prints the Phase-1 KB triples and Phase-2 lookup
trace alongside the final answer:
python examples/single_example.py \
--model-path kilian-group/KBevo-Qwen3-1.7B-GRPO \
--output-path examples/single_example_output.jsonYou can either configure resources in configs/cluster.env and use the provided wrapper:
bash scripts/submit_slurm.sh grpo-1.7bor edit the #SBATCH directives in the script and submit it directly:
sbatch scripts/grpo_variants/train_two_phase_1.7b.slurmResource options may also be passed directly to sbatch. The same approaches work for SFT and evaluation.
KBevo/
βββ configs/
β βββ accelerate/ # accelerate multi_gpu_{1,2,4,8}.yaml
β βββ artifacts.yaml # authoritative released-artifact manifest
β βββ cluster.env.example # site-specific defaults (conda env, ckpt roots, W&B, ...)
β βββ paper/ # paper configuration files (sft/grpo/eval)
βββ data/
β βββ prompts/ # database_creation.json (phase-1 prompt), lmlm_agent.json
βββ environment.yml # conda env definition
βββ examples/
β βββ mini_sft.json # 3 real SFT examples (schema + smoke)
β βββ single_example.py # one-question two-phase demo (CPU-safe --help)
β βββ expected_output.json # illustrative output schema
βββ pyproject.toml # kbevo package metadata (MIT)
βββ scripts/
β βββ eval_kbevo.sh # canonical 4-dataset two-phase eval wrapper
β βββ kbevo_hf_smoke.slurm # HF-repo-load smoke test
β βββ train_grpo.sh # canonical GRPO trainer (all sizes)
β βββ train_sft.sh # canonical SFT trainer (all sizes)
β βββ grpo_variants/ # 6 ablation slurm wrappers + README.md
βββ src/
βββ agent/ # TwoPhaseAgent, LMLMAgent (canonical inference)
βββ data/ # HotpotQA, MuSiQue, 2WikiMultiHopQA, PopQA loaders
βββ eval/ # evaluate.py (F1/EM), metrics.py
βββ eval_multihop.py # top-level eval driver
βββ grpo_train.py # GRPO trainer entrypoint (called by train_grpo.sh)
βββ llm/ # vLLM / HF backends
βββ multi_lmlm/ # DatabaseManager, TopKRetriever, prompts, constants
βββ reward_func.py # F1 / F1-format outcome rewards
βββ sft_train.py # SFT trainer entrypoint (called by train_sft.sh)
βββ tools/ # merge_shard_results, merge_unified_dbs (standalone CLIs)
βββ trainer/ # lmlm_basetrainer (two-phase GRPO advantage)
The paper trains Qwen3-1.7B and Qwen3-4B on 6k Gemini-generated HotpotQA two-phase
trajectories for 3 epochs, effective batch 48, lr 5e-5. The SFT dataset is released as
kilian-group/KBevo-SFT-hotpotqa-6k
(status: see configs/artifacts.yaml).
# Download the SFT dataset once into the repo's data/ dir (git-ignored),
# then point --dataset_path at the resulting trajectories.json.
huggingface-cli download kilian-group/KBevo-SFT-hotpotqa-6k \
--repo-type dataset --local-dir data/sft
bash scripts/train_sft.sh \
--model_size 1.7B \
--dataset_path data/sft/trajectories.jsonor set KBEVO_SFT_DATA=<path> once in configs/cluster.env and omit --dataset_path.
Add --debug to run a smoke: 1 optimizer step over the bundled
examples/mini_sft.json, isolated per-invocation output directory:
bash scripts/train_sft.sh --model_size 1.7B --debugFull hyperparameters used in the paper are recorded in
configs/paper/sft_qwen3_1.7b.yaml and
configs/paper/sft_qwen3_4b.yaml.
The paper runs two-phase GRPO from an SFT init (step 368 for 1.7B, step 735 for 4B), for 500 steps, effective batch 512, lr 5e-6. Each question fans out into K=4 Phase-1 KB rollouts and N=32 Phase-2 QA rollouts (M = N/K = 8 QAs per KB; total 36 generations/question). Retrieval uses cosine-similarity top-k=4 at threshold 0.6 over the freshly-built KB, with inverse relations enabled. The reward is token-level F1 on the final answer.
The one-line paper reproduction:
# Uses kilian-group/KBevo-Qwen3-1.7B-SFT as the init by default. Override with
# KBEVO_SFT_1_7B_CKPT=<local dir> in configs/cluster.env if you have a local ckpt.
bash scripts/grpo_variants/train_two_phase_1.7b.slurmFor 4B:
bash scripts/grpo_variants/train_two_phase_4b.slurmBoth launchers accept additional GRPO arguments. Add --debug to run a one-step smoke test with reduced data and a separate output directory:
bash scripts/grpo_variants/train_two_phase_1.7b.slurm --debugFull hyperparameters are in configs/paper/grpo_qwen3_1.7b.yaml
and configs/paper/grpo_qwen3_4b.yaml. Note that total_batch_size=512 is the
optimizer effective batch and is independent of the rollout structure (36 generations
per question).
The canonical evaluator runs the two-phase agent on HotpotQA (distractor), MuSiQue, 2Wiki, and PopQA and reports EM / F1 per dataset.
bash scripts/eval_kbevo.sh --model_path <local_or_hf_repo> \
[--datasets hotpotqa,musique,2wiki,popqa] \
[--num_samples 1000] \
[--save_version _mytag] \
[--output-dir ./output/main_tables]Full parameters (seed, top-k, threshold, sampling, max tokens) are recorded in
configs/paper/eval_kbevo.yaml.
--model_path may be a local checkpoint directory or a Hugging Face repo id (owner/name).
Per-dataset preds JSONs and aggregate metrics land under --output-dir (default
./output/main_tables/two_phase/<dataset>/<model_tag>/).
To score an existing preds JSON on its own, use
src/eval/evaluate.py:evaluate_file programmatically.
Paper configurations are provided under configs/paper/.
| File | Configuration |
|---|---|
sft_qwen3_1.7b.yaml |
Qwen3-1.7B SFT |
sft_qwen3_4b.yaml |
Qwen3-4B SFT |
grpo_qwen3_1.7b.yaml |
Qwen3-1.7B GRPO |
grpo_qwen3_4b.yaml |
Qwen3-4B GRPO |
eval_kbevo.yaml |
Two-phase evaluation |
The SFT and GRPO launchers load the corresponding YAML configuration by default. Pass a CLI argument to override a value for an individual run.
The launchers under scripts/grpo_variants/ reproduce the paper's main GRPO runs and ablations. They delegate to the canonical scripts/train_grpo.sh implementation.
| Variant | Purpose | Model |
|---|---|---|
train_two_phase_1.7b.slurm |
Paper main two-phase GRPO on Qwen3-1.7B | 1.7B |
train_two_phase_4b.slurm |
Paper main two-phase GRPO on Qwen3-4B | 4B |
train_one_phase.slurm |
Ablation: 1-phase SFT + static Gemini DB | 1.7B |
train_vanilla_grpo.slurm |
Ablation: vanilla GRPO advantage, N=16 | 1.7B |
train_zero_rl.slurm |
Ablation: GRPO from raw Qwen3-1.7B (no SFT) | 1.7B |
train_curriculum.slurm |
Ablation: tier-filtered curriculum over 90k | 1.7B |
@inproceedings{Noonan2026:co-evolving,
title = {Co-Evolving Structured Knowledge and Reasoning in Language Models},
author = {Ryan Thomas Noonan and Linxi Zhao and Menghan Xu and Akanksha Sarkar and Mihir Mishra and Dongyoung Go and Kilian Q. Weinberger and Yoav Artzi and Jennifer J. Sun},
booktitle = {Proceedings of the Conference on Language Modeling},
year = {2026},
url = {https://arxiv.org/abs/2608.26386}
}
