• Home
  • Features
  • Pricing
  • Docs
  • Announcements
  • Sign In

rm-hull / news-barge / 36175451397
94%

Build:
DEFAULT BRANCH: main
Ran 25 Sep 2026 06:47PM UTC
Jobs 1
Files 10
Run time 1min
Badge
Embed ▾
README BADGES
x

If you need to use a raster PNG badge, change the '.svg' to '.png' in the link

Markdown

Textile

RDoc

HTML

Rst

25 Sep 2026 06:45PM UTC coverage: 93.097% (-2.0%) from 95.142%
36175451397

push

github

web-flow
feat: add flair classifier for NER (named entity recognition) (#43)

* feat: add NER classifiers and slug renaming scripts

- Add multiple NER extraction scripts using Flair, spaCy, Google GenAI,
  and LiteRT for article entity backfilling.
- Add utility to recompute and rename article slugs based on site
config.

* refactor(scraper): integrate flair named entity recognition into
pipeline

Move Flair-based NER logic and helper functions into core modules to
automatically extract people, locations, and organizations during
article processing.

* ci: cache flair NER model in CI workflows

Add caching for the flair named entity recognition (NER) model in both
the build and scrape GitHub Actions workflows to speed up pipeline
execution times.

* style(scraper): use builtin types and clean up unused imports

- Update `NamedEntities` dataclass to use built-in `list` instead of
`List`.
- Remove unused `yaml` import from output module and sort imports in
pipeline and tests.

* style(scraper): add type hints to filter_duplicate_names

* fix(scraper): extract named entities from markdown body

Update `process_article` to extract named entities from `md_body`
instead of `description` for richer entity recognition across the
entire article content.

* ci(workflow): skip deployment on direct branch pushes

Prevent automatic deployment when code is pushed to branches, reserving
build-and-deploy for scheduled scrapes with changes and manual triggers.

* Dont automatically run on branch push

* refactor: defer expensive metadata generation

Move `article_categories` and `named_entities` calls after the already-exists & dry-run
checks to prevent unnecessary computation.

```mermaid
sequenceDiagram
    participant P as Pipeline
    P->>P: Check dry-run
    alt is dry-run or already exists
        P->>P: Return True
    else actual run
        P->>P: Generate categories
        P->>P: Extract entities
        P->>P: Write file
    end
```

* refactor: deduplicate entity... (continued)

51 of 64 new or added lines in 3 files covered. (79.69%)

499 of 536 relevant lines covered (93.1%)

0.93 hits per line

Uncovered Changes

Lines Coverage ∆ File
7
83.33
-16.67% scraper/src/classifiers.py
6
89.62
-5.0% scraper/src/text_extraction.py
Jobs
ID Job ID Ran Files Coverage
1 36175451397.1 25 Sep 2026 06:47PM UTC 10
93.1
GitHub Action Run
Source Files on build 36175451397
  • Tree
  • List 10
  • Changed 3
  • Source Changed 3
  • Coverage Changed 3
Coverage ∆ File Lines Relevant Covered Missed Hits/Line
  • Back to Repo
  • Github Actions Build #36175451397
  • cf0f0f2a on github
  • Prev Build on main (#36161871445)
  • Next Build on main (#36187053054)
  • Delete
STATUS · Troubleshooting · Open an Issue · Sales · Support · CAREERS · ENTERPRISE · START FREE TRIAL · SCHEDULE DEMO
ANNOUNCEMENTS · TWITTER · TOS & SLA · Supported CI Services · What's a CI service? · Automated Testing

© 2026 Coveralls, Inc