Research
In my tax work, a source can be on the right subject and still be wrong for the question, because it belongs to another jurisdiction or has been superseded. I study how retrieval can recognise that. The same concern with evidence and limits runs through my other work: pseudonymising confidential legal documents in the browser before they reach a model provider, land-cover detection from aerial imagery that returns calibrated claims with their evidence, and knowledge graphs that keep their sources through automated content pipelines. Next to this I fine-tune diffusion models and encoders, build a self-supervised model of market microstructure, and run interpretability experiments on open language models.
Retrieval for law and tax
Retrieval-augmented generation that represents what a document does not answer. A retriever scores how similar a source is to the question. It cannot say that the source is from the wrong jurisdiction, was superseded last year, or concerns a related but different concept. I model those failures as typed negative evidence, turn the signals into calibrated relevance probabilities, and let the type of uncertainty decide how the system answers.
| Failure | The source is | The system should |
|---|---|---|
| Scope | on topic, for another jurisdiction or entity | exclude it |
| Sibling | about a related but different concept | disambiguate |
| Temporal | correct once, now superseded | resolve the version |
Working paper: What doesn't match matters more.
Using cloud LLMs on confidential documents
A browser-based pseudonymisation gateway for confidential legal and fiduciary documents in Dutch, French, German, Italian and English. Protection happens on the device before anything reaches a model provider, and the answer is restored locally. The design principle is that every optional component can only add protection, never remove it.
- Document in the browser
- Local detection and pseudonymisation
- Optional encrypted check and server model on the protected text
- Protected text to the model provider
- Answer restored locally
- A deterministic evidence engine: candidate patterns, checksum and format validation, multilingual context, negative rules for legal citations, and an explained redact, review or keep decision.
- My own multilingual in-browser NER model (a token classifier exported to ONNX), trained in rounds and compared against public models on the same documents.
- A server-side model that reads only the already-protected text, and an encrypted second check using homomorphic encryption (BFV), where the server holds no keys.
- Evaluation at the egress: scoring the bytes that actually leave the device, with re-identification and utility tracks, against public anonymisation benchmarks.
Land cover from aerial imagery, Switzerland
Forest, grassland, shrub, bare rock and soil from high-resolution aerial imagery, built as a service whose output is structured. Every result is a claim: a class from an ontology, a calibrated probability, the evidence behind it and an observability band. The design objective is that the service declines claims it cannot support.
- Hand-built texture features combined with terrain and canopy-height data from national open data (elevation model, LiDAR surface model).
- Per-class calibration with reliability checks, on sites held out from training.
- Schema-validated structured output through an API, including a prototype endpoint that estimates non-eligible area inside farm parcels for subsidy checks.
- Vision-language models used as baselines and labelling tools, compared with the feature-based classifier.
- Forest stand 0.942
- Grassland 0.891
- Bare soil 0.842
- Bare rock 0.810
- Dwarf shrub 0.758
Five land-cover classes; 500 training sites and 250 calibration sites drawn from the national survey grid, scored on sites held out from training. Transfer to other countries' imagery is poor, and I report that too.
- Bare soil, raw 0.263
- Bare soil, calibrated 0.007
- Dwarf shrub, raw 0.310
- Dwarf shrub, calibrated 0.003
Lower is better. The raw scores of the rarest classes are badly overconfident; calibration removes most of that.
Species-accurate image generation
Image models draw a plausible bird or gentian, not necessarily the right species. I am building a generator that renders the correct species, and a verifier that can check whether an image shows it, for animals and plants across Europe.
- A species database of more than 230,000 European animal and plant taxa, merged from dozens of open taxonomic and trait sources (national checklists, Fauna Europaea, GBIF, iNaturalist, Tree of Life) with per-claim provenance, plus hundreds of curated trait briefs for the hardest look-alikes.
- Trait-conditioned prompting, LoRA, DreamBooth and textual-inversion fine-tuning of FLUX-class and Z-Image models, and ControlNet arms, on licence-filtered photographs.
- A multi-juror gate: BioCLIP nearest-centroid scoring against real photographs plus vision-language jurors, with controls built from confusable congeners.
- Species name only 0.222 12/54
- Name + named traits 0.519 28/54
- Name + shuffled traits (control) 0.056 3/54
Nine look-alike gentian species, 54 renders per prompt type with identical seeds, scored by BioCLIP against 287 real photographs. The shuffled-trait control separates information from decoration.
Knowledge graphs and automated content pipelines
A layered knowledge graph in which every claim carries its sources and evidence. Automated pipelines build it, with LLM steps bound by schemas and contracts and checked by tests, and turn it into narrated walking tours.
- Sources
- Evidence
- Claims and entities
- Discovery and importance
- Editorial layers
- Published content
Routes and maps
Hiking and walking routes synthesised from crowd-sourced GPS traces and cross-checked against OpenStreetMap: a consensus centreline, snapping to the path network with surface and difficulty attributes, stage splitting, and 3D flyover rendering. Pedestrian routing runs on OpenStreetMap graphs, with one routing engine per country.
Models and tooling
- Encoders: token classifiers for multilingual NER, and an XLM-RoBERTa classifier for detecting machine-written text in four languages.
- Local language models: benchmarking and serving open models on a single GPU, with speculative-decoding sweeps.
- Agent tooling: MCP servers, cost-aware model routing behind a privacy gate, and evaluation harnesses.
Markets
A self-supervised neural encoder on trade and quote data (volume clock, multi-channel microstructure windows) with frozen linear readouts, tested against standard volatility models (HAR-RV and variants). A companion study asks whether market histories are better grouped by what followed them than by what they look like, using contrastive embeddings and JEPA-style variants.
Earlier: a long-only equity strategy on QuantConnect. It ranks S&P 500 stocks by a crash-risk measure (down-to-up volatility, after Chen, Hong and Stein, 2001), tracked weekly over several time frames together with its speed and acceleration. Entries come from a signal ladder, and a scored daily check handles exits. Backtested from 2019 on a cash account.
Interpretability
Logit-lens, lens-vector, ablation and coefficient-swap experiments on a small open language model, replicating published results.