browsertrix

Author	SHA1	Message	Date
Ilya Kreymer	9553115bbe	helm chart tweaks: (#1067 ) * helm chart tweaks: - lower mem requirements for backend and crawler - disable cors in ingress to pass through cors headers from backend - crawler statefulset: use ordered instead of parallel scaling policy to avoid single crawl taking up all crawling capacity quickly	2023-08-14 16:43:12 -07:00
sua yoo	ffd0e525d9	Webpack config improvements (#1063 ) - Upgrades webpack and webpack-dev-server for bugfixes and performance updates - Removes unnecessary file watching - Enables persistent build cache in dev - Switches to faster dev source map	2023-08-11 13:16:24 -07:00
Ilya Kreymer	d93ddaf620	bump version to 1.6.1	2023-08-11 12:50:41 -07:00
Ilya Kreymer	35ab6d6df6	bump to 1.6.0!	2023-08-09 15:40:27 -07:00
Ilya Kreymer	8ea3dd5dae	terminology tweaks in frontend: (part of #922 ) (#1062 ) * terminology tweaks in frontend: (part of #922) - use 'crawl workflow' instead of 'workflow' where possible - use 'replay' instead of 'replay crawl' - localization: rerun string extraction / processing - "Review Config" → "Review Settings" - "Workflow" → "Crawl Workflow" in error message --------- Co-authored-by: Henry Wilkinson <henry@wilkinson.graphics>	2023-08-09 15:38:58 -07:00
sua yoo	37733483d5	Standardize archived item filtering, sorting and labels (#1054 ) Frontend: - Renames list view to "All Archived Items" - Refactors fetches to use single all-crawls endpoints - Removes search by config ID for more search parity with uploads - Adds sort by size - Refactors property and method names to replace crawl* - Replaces remaining references to "crawl" in copy with "item"' - Rename Upload Archive button to Upload WACZ - Fix focusout in item menu so menus close Backend: - Filter search values by type as well - Only get list of cids for crawls in search values - Don't list crawl/workflow ids in search values --------- Co-authored-by: Tessa Walsh <tessa@bitarchivist.net>	2023-08-09 12:13:55 -07:00
Ilya Kreymer	7a8f370bc2	bump version to 1.6.0-beta.4 for testing	2023-08-09 12:09:37 -07:00
Ilya Kreymer	38f67a6cc0	Optimize Frontend Image Build on CI (#1057 ) * Always run yarn only on build platform with --platform=$BUILDPLATFORM * Remove optional dependencies (playwright + chromium) from build with --ignore-optional and move some devDependencies to be optional * Disable husky pre-commit hook checks on frontend Co-authored-by: sua yoo <sua@suayoo.com>	2023-08-09 12:06:20 -07:00
sua yoo	b494070e43	Collection share dialog + copy updates (#1056 ) - Always shows primary "Share" action button in Collection detail page. - Enables toggling shareable status and share info from dialog. Difference from mockups: I made the "Done" button neutral do differentiate from our submit action buttons in the dialog, since toggling will apply changes immediately. - Menu item: "Go to Public View"/"Go to Shareable View" -> "Visit Shareable URL". - Toggle label: "Make Collection Shareable" -> "Collection is Shareable". - Additional dialog copy: adds "This collection can be viewed by anyone with the link." under "Link to Share" and "Share this collection by embedding it into an existing webpage." under "Embed Collection". - Moves share status icon to its own column in list view. - Adds new syntax-highlighted code component that supports js and html. Co-authored-by: Henry Wilkinson <henry@wilkinson.graphics>	2023-08-09 10:12:46 -07:00
Ilya Kreymer	de3e5907a7	backend: crawlout: include raw crawnconfig in api details, fixes #1030 (#1055 )	2023-08-09 08:46:42 -07:00
Ilya Kreymer	8d0a4f2ca9	fix public collections endpoint returning 404 when not public (#1052 ) tests: add tests for public collections endpoint when collection is public and when not	2023-08-04 13:29:13 -04:00
Tessa Walsh	7ff57ce6b5	Backend: standardize search values, filters, and sorting for archived items (#1039 ) - all-crawls list endpoint filters now conform to 'Standardize list controls for archived items #1025' and URL decode values before passing them in - Uploads list endpoint now includes all all-crawls filters relevant to uploads - An all-crawls/search-values endpoint is added to support searching across all archived item types - Crawl configuration names are now copied to the crawl when the crawl is created, and crawl names and descriptions are now editable via the backend API (note: this will require frontend changes as well to make them editable via the UI) - Migration added to copy existing config names for active configs into their associated crawls. This migration has been tested in a local deployment - New statuses generate-wacz, uploading-wacz, and pending-wait are added when relevant to tests to ensure that they pass - Tests coverage added for all new all-crawls endpoints, filters, and sort values	2023-08-04 09:56:52 -07:00
Anish Lakhwara	9236a07800	fix: run `yarn format` in frontend dir (#1043 )	2023-08-03 19:12:48 -07:00
Ilya Kreymer	362afa47bd	Support for Public / Shareable Collections (#1038 ) * collections: support toggling collections public/private, viewable via RWP - backend: add 'public' to collection model, support patching to update - backend: add .../collections/<id>/public/replay.json for public access - backend: add CORS handling for public endpoint - frontend: support 'make shareable / make private' dropdown actions on collection detail + collection list views - frontend: show shareable / private icons by collection name on detail + list views - frontend: link to replayweb.page for standalone browsing - frontend: add embed code popup when a collection is shareable - refer to public collections as 'shareable' for now --------- Co-authored-by: Henry Wilkinson <henry@wilkinson.graphics>	2023-08-03 19:11:01 -07:00
sua yoo	62d3399223	Add info bar to Collection detail view (#1036 ) - Adds Collection info bar to detail view - Update "Web Captures" -> "Archived Items" - Updates Collection list columns to match - Refactors `btrix-desc-list` and usage in `workflow-details` to reuse horizontal info bar component	2023-08-03 16:58:56 -07:00
Anish Lakhwara	af09d56ef6	Merge pull request #1035 from webrecorder/backend-init feat: Display waiting message while backend is initializing	2023-08-02 17:39:47 -07:00
Anish Lakhwara	fa58e77167	fix: remove strange character?	2023-08-02 17:34:09 -07:00
Anish Lakhwara	5ed2faaecc	fix: need to use `window.timeOut` to get a timerId back	2023-08-02 17:31:01 -07:00
Anish Lakhwara	6ecfd8ec24	fix: timerId not timeoutId	2023-08-02 17:28:07 -07:00
Anish Lakhwara	3985cf014e	fix: clear timeout on disconnect callback	2023-08-02 17:26:26 -07:00
Anish Lakhwara	196b26c60e	fix: center text	2023-08-02 17:21:36 -07:00
Anish Lakhwara	f1d91e3bf9	fix: add styling	2023-08-02 17:18:40 -07:00
Anish Lakhwara	a8bedeffb5	fix: take Sua's suggestons, less code needed	2023-08-02 17:10:45 -07:00
Anish Lakhwara	2f26fcefce	fix: make pretty & work correctly	2023-08-02 16:36:28 -07:00
Anish Lakhwara	06918c967b	feat: use html dialog instead	2023-08-02 11:37:55 -07:00
Anish Lakhwara	84a60b54e4	feat: Display waiting message while backend is initializing	2023-08-01 17:18:05 -07:00
Ilya Kreymer	45eaa0b3a3	version: bump to 1.6.0-beta.3	2023-08-01 09:48:17 -07:00
sua yoo	cc52dfd940	Sort Collections by size (#1026 ) - Adds "Size" column to Collections list view - Adds "Size" option to sort dropdown	2023-08-01 09:47:47 -07:00
Anish Lakhwara	32428f4d93	fix: usr/bin/env bash interpreter for btrix (#1028 )	2023-08-01 09:28:56 -07:00
sua yoo	54e2b2c703	List web captures in Collection (#1024 ) - Adds tab for "Web Captures" in Collection detail view - Move Collection description under Replay section - Fixes app reloading when clicking into a Collection - Standardizes Web Capture list headers from "Finished -> "Created Date"	2023-08-01 09:14:27 -07:00
Ilya Kreymer	06cf9c7cc3	add crawl ending states: 'generate-wacz', 'uploading-wacz', 'pending-wait' that occur after a crawl is finished or is being stopped (#1022 ) operator: ensure transitions from each of these states is supported, including to 'waiting_capacity' add extra check on stopping to avoid transitioning back to a running state after crawl is finished ui: add states to UI display, localization, add as active states fixes #263	2023-08-01 00:15:59 -07:00
Anish Lakhwara	d848502f84	fix(build): add DOCKER_BUILDKIT=1 to frontend Dockerfile to better support older versions of Docker (#1021 )	2023-08-01 00:15:03 -07:00
Ilya Kreymer	7ea6d76f10	Resource Constraints Cleanup: (fixes #895 ) (#1019 ) * resource constraints: (fixes #895) - for cpu, only set cpu requests - for memory, set mem requests == mem limits - add missing resource constraints for minio and scheduled job - for crawler, set mem and cpu constraints per browser, scale based on browser instances per crawler - add comments in values.yaml for crawler values being multiplied - default values: bump crawler to 650 millicpu per browser instance just in case cleanup: remove unused entries from main backend configmap	2023-08-01 00:11:16 -07:00
Anish Lakhwara	d8502da885	fix(build): use `/usr/bin/env bash` instead of `/bin/bash` (#1020 ) * fix: add to various other shell scripts	2023-07-28 21:50:04 -07:00
Ilya Kreymer	c76dd10928	chart: always pull latest crawler image - since default image is pointing to webrecorder/browsertrix-crawler:latest, makes sense to always pull latest (#1018 )	2023-07-27 12:41:41 -07:00
Anish Lakhwara	a347f61973	ci: password check: fix: don't break on ScannerError (#1017 )	2023-07-27 07:19:27 -07:00
Vinzenz Sinapius	5807507f29	Add proxy settings for crawler and profilebrowser (#997 )	2023-07-26 16:11:10 -07:00
sua yoo	7069b33646	Show only running crawls in superadmin view (#1015 ) - Show separate crawls list for admin view, fixes #1010	2023-07-26 15:48:20 -07:00
Ilya Kreymer	6506965d98	Streaming Download for Collections (#1012 ) * support streaming download of collections (part of #927) - WACZ zip created on the fly using stream-zip - add 'Download Collection' option to collection detail and list - after editing collection, return to collection view - tests: add test for streaming download, ensure WACZ files + datapackage present, STORE compression used --------- Co-authored-by: sua yoo <sua@suayoo.com>	2023-07-26 15:42:17 -07:00
Anish Lakhwara	6062042fae	feat: create DO registry if it doesn't exist (#947 ) - if use_do_registry is enabled and registry doesn't exist, create it	2023-07-26 15:41:03 -07:00
Anish Lakhwara	4c1465d94b	feat: ansible DO teardown (#950 ) * feat: ansible DO teardown * fix(DO): idempotency issues in ansible teardown * chore(DO): remove unused code * docs(ansible): mention teardown in the docs * fix: pass ansible-lint * fix: point database backup upload to the correct location in DO space	2023-07-26 15:38:59 -07:00
Anish Lakhwara	b5a9c42df1	feat: add pre-commit to check we don't have real passwords in yml files (#990 ) * feat: use existing pre-commit framework * feat(ci): add github action for password_check * feat: add some simple tests to password_check.py * fix: set `backend_password_secret` in default values.yaml to an allowed password	2023-07-26 13:29:37 -07:00
Tessa Walsh	c21153255a	Rename notes to description in frontend and backend (#1011 ) - Rename crawl notes to description - Add migration renaming notes -> description - Stop inheriting workflow description in crawl - Update frontend to replace crawl/upload notes with description - Remove setting of config description from crawl list - Adjust tests for changes	2023-07-26 13:00:04 -07:00
Ilya Kreymer	4bea7565bc	load handling: scale up redis only when crawler pods running (#1009 ) Operator: Modified init behavior to only load redis when at least one crawler pod available: - waits for at least one crawler pod to be available before starting redis pod, to avoid situation where many crawler pods are in pending mode, but redis pods are still running. - redis statefulset starts at scale of 0 - once crawler pod becomes available, redis sts is scaled to 1 (via `initRedis==true` status) - crawl remains in 'starting' or 'waiting_capacity' state until pod becomes available without redis pod running - set to 'running' state only after redis and at least one crawler pod is available - if no crawler pods available after running, or, if stuck in starting for >60 seconds, switch to 'waiting_capacity' state - when switching to 'waiting_capacity', also scale down redis to 0, wait for crawler pod to become available, only then scale up redis to 1, and get back to 'running' other tweaks: - add new status field 'initRedis', default to false, not displayed - crawler pod: consider 'ContainerCreating' state as available, as container will not be blocked by resource limits - add a resync after 3 seconds when waiting for crawler pod or redis pod to become available, configurable via 'operator_fast_resync_secs' - set_state: if not updating state, ensure state reflects actual value in db	2023-07-26 08:40:05 -07:00
Tessa Walsh	608a744aaf	Add migration to replace None with 0 for configmap CRAWL_TIMEOUT (#1008 )	2023-07-24 15:49:26 -04:00
Tessa Walsh	fcd48b1831	Add totalSize to collections and make it sortable in list endpoint (#1001 ) * Precompute collection.totalSize and make sortable * Add migration to recompute collection data with totalSize	2023-07-24 13:12:23 -04:00
sua yoo	75b011f951	Upload WACZ via UI (#992 ) - Users can now upload .WACZ archives from the "Archived Data" page. - Can specify name, description, tags and collection(s) to add upload to - Show progress of upload - Support canceling upload	2023-07-21 16:45:52 +02:00
Tessa Walsh	9f32aa697b	Add collections and tags to upload API endpoints (#993 ) * Add collections and tags to uploads * Fix order of deletion check test * Re-add tags to UploadedCrawl model after rebase * Fix Users model heading	2023-07-21 16:44:56 +02:00
Tessa Walsh	4014d98243	Move pydantic models to separate module + refactor crawl response endpoints to be consistent (#983 ) * Move all pydantic models to models.py to avoid circular dependencies * Include automated crawl details in all-crawls GET endpoints - ensure /all-crawls endpoint resolves names / firstSeed data same as /crawls endpoint for crawls to ensure consistent frontend display. fields added in get and list all-crawl endpoints for automated crawls only: - cid - name - description - firstSeed - seedCount - profileName * Add automated crawl fields to list all-crawls test * Uncomment mongo readinessProbe * cleanup CrawlOutWithResources: - remove 'files' from output model, only resources should be returned - add _files_to_resources() to simplify computing presigned 'resources' from raw 'files' - update upload tests to be more consistent, 'files' never present, 'errors' always none --------- Co-authored-by: Ilya Kreymer <ikreymer@gmail.com>	2023-07-20 13:05:33 +02:00
Tessa Walsh	577416024b	Fix pull_request syntax in ansible lint GH Action (#995 ) * Fix pull_request syntax in ansible lint GH Action * Only lint Digital Ocean playbook for now * fix: pass ansible lint --------- Co-authored-by: Anish Lakhwara <anish+git@lakhwara.com>	2023-07-20 12:13:52 +02:00

1 2 3 4 5 ...

741 Commits