linkerd2

Commit Graph

Author	SHA1	Message	Date
Risha Mars	136b9cc7c1	Add linkerd check flag to run data plane checks (#1528 ) Adds a --proxy flag to the linkerd check CLI command which will run to-be-implemented data plane checks	2018-08-28 10:16:24 -07:00
Risha Mars	fff09c5d06	Only tap pods that are meshed (#1535 ) Previously, we would tap any resource's pods, regardless of whether the pods were meshed or not. We can't actually tap non-meshed pods, so I'm adding a check that will filter out non-meshed pods from the pods that tap watches. Previous behaviour: When attempting to hang a non meshed pod, it would establish a watch on the pods, but then never return any results. In the CLI you could just cancel it with Ctrl-C. In the web, clicking Stop would send a WebSocket.close(1000) but wouldn't actually close the connection... Behaviour after change : If no pods under the specified resource are meshed, it'll return an error of no pods being found to tap	2018-08-28 09:59:52 -07:00
Risha Mars	27e52a6cc0	Add ReadinessProbe and LivenessProbe to injected proxy containers (#1530 ) Adds basic probes to the linkerd-proxy containers injected by linkerd inject. - Currently the Readiness and Liveness probes are configured to be the same. - I haven't supplied a periodSeconds, but the default is 10. - I also set the initialDelaySeconds to 10, but that might be a bit high. https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-probes/	2018-08-27 11:55:17 -07:00
Kevin Lingerfelt	4450a7536d	Add --wait flag for CLI check and dashboard commands (#1503 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-22 12:56:42 -07:00
Kevin Lingerfelt	49f6c4c770	Refactor healthcheck init and observe setup (#1502 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-22 12:30:45 -07:00
Kevin Lingerfelt	53cd3b50d5	Add --pre flag for linkerd check command (#1497 ) * Add --pre flag for linkerd check command * Small adjustments to check help text Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-20 17:09:43 -07:00
Kevin Lingerfelt	e97be1f5da	Move all healthcheck-related code to pkg/healthcheck (#1492 ) * Move all healthcheck-related code to pkg/healthcheck * Fix failed check formatting * Better version check wording Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-20 16:50:22 -07:00
Eliza Weisman	b8434d60d4	Add resource metadata to Tap CLI output (#1437 ) Closes #1170. This branch adds a `-o wide` (or `--output wide`) flag to the Tap CLI. Passing this flag adds `src_res` and `dst_res` elements to the Tap output, as described in #1170. These use the metadata labels in the tap event to describe what Kubernetes resource the source and destination peers belong to, based on what resource type is being tapped, and fall back to pods if either peer is not a member of the specified resource type. In addition, when the resource type is not `namespace`, `src_ns` and `dst_ns` elements are added, which show what namespaces the the source and destination peers are in. For peers which are not in the Kubernetes cluster, none of these labels are displayed. The source metadata added in #1434 is used to populate the `src_res` and `src_ns` fields. Also, this branch includes some refactoring to how tap output is formatted. Signed-off-by: Eliza Weisman <eliza@buoyant.io>	2018-08-20 14:25:26 -07:00
Kevin Lingerfelt	7c07ba0d53	Upgrade to dep 0.5.0, go 1.10.3 (#1479 ) * Upgrade to dep 0.5.0, go 1.10.3 * Remove existing dep binary if it's the wrong version * Add version in filename of dep binary to prevent version conflicts Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-17 16:04:50 -07:00
Alex Leong	094a375015	[RFC] linkerd top (#1435 ) This an initial implementation of the `linkerd top` command. This command launches an ncurses style tabular view of current requests (using data from tap). Most of the command line arguments are the same as tap and allow selecting the resource to inspect and filtering which requests to view. Fixes #1283 Signed-off-by: Alex Leong <alex@buoyant.io>	2018-08-15 18:10:23 -07:00
Kevin Lingerfelt	00a0572098	Better CLI error messages when control plane is unavailable (#1428 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-09 15:40:41 -07:00
Eliza Weisman	9d8f58cb16	Add additional validation for stat command-line arguments (#1415 ) Closes #776. This branch adds the following validation to the `linkerd stat` command: * The `--to` and `--from` flags are now mutually exclusive * The `--to-namespace` and `--from-namespace` commands are also mutually exclusive. * The `namespace` resource type conflicts with the `--namespace`, `--to-namespace`, and `--from-namespace` flags. Examples: ``` $ bin/go-run cli/main.go stat deploy --to deploy/foo --from deploy/bar Error: --to and --from flags are mutually exclusive Usage: linkerd stat [flags] (RESOURCE) ... ``` ``` $ bin/go-run cli/main.go stat deploy --to-namespace foo --from-namespace bar Error: --to-namespace and --from-namespace flags are mutually exclusive Usage: linkerd stat [flags] (RESOURCE) ... ``` ``` $ bin/go-run cli/main.go stat namespace foo --namespace bar Error: --namespace flag is incompatible with namespace resource type Usage: linkerd stat [flags] (RESOURCE) ... ``` ``` $ bin/go-run cli/main.go stat ns --to-namespace bar Error: --to-namespace flag is incompatible with namespace resource type Usage: linkerd stat [flags] (RESOURCE) ... ``` ``` $ bin/go-run cli/main.go stat namespace --from-namespace bar Error: --from-namespace flag is incompatible with namespace resource type Usage: linkerd stat [flags] (RESOURCE) ... ``` ``` $ bin/go-run cli/main.go stat ns/foo --from-namespace bar Error: --from-namespace flag is incompatible with namespace resource type Usage: linkerd stat [flags] (RESOURCE) ... ``` Signed-off-by: Eliza Weisman <eliza@buoyant.io>	2018-08-08 15:35:47 -07:00
Kevin Lingerfelt	82940990e9	Rename mailing lists, remove all remaining conduit references (#1416 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-07 17:00:55 -07:00
Kevin Lingerfelt	4845b4ec04	Restore linkerd.io/control-plane* labels (#1411 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-07 13:53:29 -07:00
Kevin Lingerfelt	e0a01c5dd8	Remove node scrape target, kubernetes grafana dashboard (#1410 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-07 13:41:38 -07:00
Kevin Lingerfelt	bd19e8aaff	Update prometheus to only scrape proxies in the same mesh (#1402 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-06 12:05:55 -07:00
Kevin Lingerfelt	f70ad7de11	Use stable version for linkerd2-proxy-api dep (#1400 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-08-03 11:59:42 -07:00
Sean McArthur	c035193313	add H2 protocol to destination addrs if managed by linkerd (#1380 ) Signed-off-by: Sean McArthur <sean@buoyant.io>	2018-08-03 10:14:30 -07:00
Eliza Weisman	01cc30d102	Increase outbound router capacity for Prometheus pod's proxy (#1358 ) Currently, when a cluster has over 100 pods injected with the Linkerd2 proxy, Prometheus metrics are not collected correctly. This is because Prometheus appears to be making more concurrent requests than its' proxy's outbound router cache can handle See issue #1322 for further details. This branch introduces a workaround for this issue, by increasing the outbound router cache capacity to 10000 routes for the Prometheus pod's proxy only. The router capacity limit of 100 active routes is primarily due to the limitation of the number of active Destination service lookups, so increasing the capacity for the Prometheus pod specifically is probably okay, as the scrape requests are made to IP addresses directly and therefore will not cause service discovery lookups. This change was originally implemented and tested in @siggy's PR #1228. I've rebased his branch onto the current `master`, and updated the code to reflect the project name change. Signed-off-by: Eliza Weisman <eliza@buoyant.io> Co-authored-by: Andrew Seigner <siggy@buoyant.io>	2018-08-02 16:44:11 -07:00
Ivan Sim	eb04217a12	Update inject cmd to read from folder (#1377 ) This change is a simplified implementation of the Builder.Path() and Visitor().ExpandPathsToFileVisitors() functions used by kubectl to parse files and directories. The filepath.Walk() function is used to recursively traverse directories. Every .yaml or .json resource file in the directory is read into its own io.Reader. All the readers are then passed to the YAMLDecoder in the InjectYAML() function. Fixes #1376 Signed-off-by: ihcsim <ihcsim@gmail.com>	2018-08-01 17:12:00 -07:00
Kevin Lingerfelt	8fe9e53f67	Remove remaining conduit references in codebase (#1381 ) * Remove remaining conduit references in codebase * Shorten emojivoto config url Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-31 11:19:34 -07:00
Kevin Lingerfelt	c362d5e114	Update k8s.io dependencies to 1.11.1 (#1369 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-27 15:23:03 -07:00
Kevin Lingerfelt	51848230a0	Send glog logs to stderr by default (#1367 ) * Send glog logs to stderr by default * Factor out more shared flag parsing code Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-25 12:59:24 -07:00
Risha Mars	ec3c861743	Enable Tap from the Web UI (#1356 ) Adds a tap endpoint in the web api that communicates with the dashboard via websockets. I've moved a bunch of code from the cli tap.go into utils so that the code can be shared between web and CLI. I think we should consider making the display more suited to web, but in the short term, reusing the CLI's rendering of tap events works. Adds a Tap page in the Web UI that you can use to make tap requests. The form currently only allows you to enter a resource and namespace, other filters coming in a follow-up branch.	2018-07-24 14:23:42 -04:00
Kevin Lingerfelt	4b9700933a	Update prometheus labels to match k8s resource names (#1355 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-23 15:45:05 -07:00
Brian Smith	a98bfb1ca7	Rename `ca-bundle-distributor` to `ca`. (#1340 ) `ca-bundle-distributor` described the original role of the program but `ca` ("Certificate Authority") better describes its current role. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-07-17 14:10:40 -10:00
Brian Smith	1b38310019	Remove executable bit from non-executable files. (#1335 ) These files were created with the executable bit set accidentally due to the way my network file system setup was configured. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-07-16 13:55:52 -10:00
Brian Smith	0fcfd2bffb	Stop using `installsuffix` when building Go code. (#1327 ) * Stop using `installsuffix` when building Go code. See https://plus.google.com/117192131596509381660/posts/eNnNePihYnK. `-installsuffix cgo` isn't necessary as of Go 1.10 (where build caching changed substantially) and it probably wasn't necessary earlier. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-07-16 13:48:50 -10:00
Franziska von der Goltz	c7ac072acc	update grafana dashboards: conduit to linkerd (#1320 ) * update grafana dashboards to remove conduit reference and replace with linkerd instances * update test install fixtures to reflect changes Fixes: #1315 Signed-off-by: Franziska von der Goltz <franziska@vdgoltz.eu>	2018-07-16 13:05:01 -07:00
Kevin Lingerfelt	e5cce1abaf	Rename CLI from conduit to linkerd (#1312 ) * Rename CLI binary * Update integration tests for new binary name * Rename --conduit-namespace flag, change default ns * Rename occurrences of conduit in rest of CLI * Rename inject and install components * Remove conduit occurrences in docker files * Additional miscellaneous cleanup * Move protobuf definitions to linkerd2 package * Rename conduit.io labels to use linkerd.io * Rename conduit-managed segment to linkerd-managed * Fix conduit references in web project Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-12 17:14:07 -07:00
Andrew Seigner	e18fa48135	Name ClusterRole objects to be namespace-specific (#1295 ) The control-plane's `ClusterRole` and `ClusterRoleBinding` objects are global. Because their names did not vary across multiple control-plane deployments, it prevented multiple control-planes from coexisting (when RBAC is enabled). Modify the `ClusterRole` and `ClusterRoleBinding` objects to include the control-plane's namespace in their names. Also modify the integration test to first install two control-planes, and then perform its full suite of tests, to prevent regression. Fixes #1292. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-07-10 16:21:20 -07:00
Oliver Gould	941cad4a9c	Migrate build infrastructure to linkerd2 (#1298 ) This PR begins to migrate Conduit to Linkerd2: * The proxy has been completely removed from this repo, and is now located at github.com/linkerd/linkerd2-proxy. * A `Dockerfile-proxy` has been added to fetch the most-recently published proxy binary from build.l5d.io. * Proxy-specific protobuf bindings have been moved to github.com/linkerd/linkerd2-proxy-api. * All docker images now use the gcr.io/linkerd-io registry. * `inject` now uses `LINKERD2_PROXY_` environment variables * Go paths have been updated to reflect the new (future) repo location.	2018-07-09 15:38:38 -07:00
Kevin Lingerfelt	fd1aecfa63	Unhide --tls flag in conduit CLI (#1278 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-05 15:49:19 -07:00
Kevin Lingerfelt	693acdbf26	Update ListPods endpoint to return all pod owner types (#1275 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-05 15:14:16 -07:00
Kevin Lingerfelt	f0ba8f3ee8	Fix owner types in TLS identity strings (#1257 ) * Fix owner types in TLS identity strings * Update documentation on TLSIdentity struct Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-03 14:20:24 -07:00
Risha Mars	83b982b25a	Change CLI and web TLS indicators from Secured to TLS (#1247 ) Previously, we had "Secured" columns in the web and CLI for the percentage of traffic that is TLSed. Change this to "TLS"	2018-07-03 10:51:38 -07:00
Brian Smith	252a8d39d3	Generate an ephemeral CA at startup that distributes TLS credentials (#1245 ) Create a ephemeral, in-memory TLS certificate authority and integrate it into the certificate distributor. Remove the re-creation of deleted ConfigMaps; this will be added back later in #1248. Signed-off-by: Brian Smith brian@briansmith.org Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-07-02 18:09:31 -10:00
Oliver Gould	20276b106e	tap: Support `tls` labeling (#1244 ) The proxy's metrics are instrumented with a `tls` label that describes the state of TLS for each connection and associated messges. This same level of detail is useful to get in `tap` output as well. This change updates Tap in the following ways: * `TapEvent` protobuf updated: * Added `source_meta` field including source labels * `proxy_direction` enum indicates which proxy server was used. * The proxy adds a `tls` label to both source and destination meta indicating the state of each peer's connection * The CLI uses the `proxy_direction` field to determine which `tls` label should be rendered.	2018-07-02 17:19:20 -07:00
Kevin Lingerfelt	a685dba873	Use parent name instead of pod name in identity string (#1236 ) * Use parent name instead of pod name in identity string * Update protobuf comment Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-29 14:28:13 -07:00
Brian Smith	f989c56127	Proxy: Skip TLS for control plane loopback connections. (#1229 ) If the controller address has a loopback host then don't use TLS to connect to it. TLS isn't needed for security in that case. In mormal configurations the proxy isn't terminating TLS for loopback connections anyway. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-06-28 17:24:09 -10:00
Risha Mars	5ed7fc563c	Add controller component pod uptimes to the ServiceMesh page (#1205 ) - Return pod uptimes from the GetPods endpoint - Adds filtering by namespace to api.GetPods - Adds a --namespace filter to conduit get pods - Adds pod uptimes to the controller component toolitps on the ServiceMesh page - Moves the ServiceMesh page back to using /api/pods	2018-06-28 15:42:00 -07:00
Risha Mars	68586fe697	Add the ability to query stats by authority (#1181 ) Adds the ability to query by a new non-kubernetes resource type, "authorities", in the StatSummary api. This includes an extensive refactor of stat_summary.go to deal with non-kubernetes resource types. - Add documentation to Resource in the public api so we can use it for authority - Handle non-k8s resource requests in the StatSummary endpoint - Rewrite stat summary fetching and parsing to handle non-k8s resources - keys stat summary metric handling by Resource instead of a generated string - Adds authority to the CLI - Adds /authorities to the Web UI - Adds some more stat integration and unit tests	2018-06-28 14:31:44 -07:00
Kevin Lingerfelt	ef9c890505	Fix issue with injected resource name, add test (#1226 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-28 10:23:38 -10:00
Oliver Gould	9f274526d6	cli: tap: Use safe accessors (#1224 ) The `tap` command is prone to panic due to use of `nil` values. This is because we don't use the safe `Get*()` field accessors provided by protobuf. This change fixes several unsafe field access paths. Fixes #47	2018-06-28 11:10:56 -07:00
Thomas Rampelberg	fafce1b8b3	Add important comment back (#1219 )	2018-06-28 08:18:52 -07:00
Brian Smith	cca8e7077d	Add TLS support to `conduit inject`. (#1220 ) * Add TLS support to `conduit inject`. Add the settings needed to enable TLs when `--tls=optional` is passed on the commend line. Later the requirement to add `--tls` will be removed. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-06-27 16:04:07 -10:00
Thomas Rampelberg	97868f654f	Add Pod to injectable types (#1213 ) * Add Pod to injectable types * Remove the pod label for pods	2018-06-27 14:37:05 -07:00
Kevin Lingerfelt	1f1968ad4d	Add --registry flag support for inject command (#1188 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-22 12:52:42 -07:00
Kevin Lingerfelt	5cf8ab00df	Switch to multi-value --tls flag, add to inject (#1182 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-21 15:52:14 -07:00
Kevin Lingerfelt	af85d1714f	Add probes and log termination policy for distributor (#1178 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-21 14:02:41 -07:00
Kevin Lingerfelt	12f869e7fc	Add CA certificate bundle distributor to conduit install (#675 ) * Add CA certificate bundle distributor to conduit install * Update ca-distributor to use shared informers * Only install CA distributor when --enable-tls flag is set * Only copy CA bundle into namespaces where inject pods have the same controller * Update API config to only watch pods and configmaps * Address review feedback Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-21 13:12:21 -07:00
Kevin Lingerfelt	e80356de34	Upgrade prometheus to v2.3.1 (#1174 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-21 11:02:21 -07:00
Kevin Lingerfelt	682b0274b5	Add controller admin servers and readiness probes (#1168 ) * Add controller admin servers and readiness probes * Tweak readiness probes to be more sane * Refactor based on review feedback Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-20 17:32:44 -07:00
Risha Mars	0ff1bb4ad8	Don't allow stat requests for named resources in --all-namespaces (#1163 ) Don't allow the CLI or Web UI to request named resources if --all-namespaces is used. This follows kubectl, which also does not allow requesting named resources over all namespaces. This PR also updates the Web API's behaviour to be in line with the CLI's. Both will now default to the default namespace if no namespace is specified.	2018-06-20 12:59:31 -07:00
Risha Mars	46c99febf2	Don't panic on stats that aren't included in StatAllResourceTypes (#1154 ) Problem `conduit stat` would cause a panic for any resource that wasn't in the list of StatAllResourceTypes This bug was introduced by https://github.com/runconduit/conduit/pull/1088/files Solution Fix writeStatsToBuffer to not depend on what resources are in StatAllResourceTypes Also adds a unit test and integration test for `conduit stat ns`	2018-06-19 17:00:16 -07:00
Andrew Seigner	0b9e7ff7df	Enable get for nodes/proxy for Prometheus RBAC (#1142 ) The `kubernetes-nodes-cadvisor` Prometheus queries node-level data via the Kubernetes API server. In some configurations of Kubernetes, namely minikube and at least one baremetal kubespray cluster, this API call requires the `get` verb on the `nodes/proxy` resource. Enable `get` for `nodes/proxy` for the `conduit-prometheus` service account. Fixes #912 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-06-18 17:49:23 +01:00
Risha Mars	e2c2f19d2c	Propagate errors in conduit containers to the api (#1117 ) - It would be nice to display container errors in the UI. This PR gets the pod's container statuses and returns them in the public api - Also add a terminationMessagePolicy to conduit's inject so that we can capture the proxy's error messages if it terminates	2018-06-14 16:22:31 -07:00
Thomas Rampelberg	516807bde6	Add readiness/liveness checks for third party components (#1121 ) * Add readiness/liveness checks for third party components Any possible issues with the third party control plane components can wedge the services. Take the best practices for prometheus/grafana and add them to our template. See #1116 * Update test fixtures for new output	2018-06-14 13:01:13 -07:00
Kevin Lingerfelt	9f1df963e9	Move controller/util and web/util packages to pkg (#1109 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-13 11:25:56 -07:00
Kevin Lingerfelt	b6d429e80d	dst svc: use shared informer instead of custom endpoints informer (#1079 ) * Update destination service ot use shared informer instead of custom endpoints informer * Add additional tests for dst svc endpoints watcher * Remove service ports when all listeners unsubscribed * Update go deps Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-13 11:11:57 -07:00
Kevin Lingerfelt	6e66f6d662	Rename Lister to API and expose informers as well as listers (#1072 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-12 10:27:55 -07:00
Risha Mars	7d4c4aa290	CLI: print resources in the same order every time stat all is run (#1088 ) Previously, in conduit stat all we would just print the map of stat results, which resulted in the order in which stats were displayed varying between prints. Fix: Define an array, k8s.StatAllResourceTypes and use the order in this array to print the map; ensuring a consistent print order every time the command is run.	2018-06-08 15:02:17 -07:00
Kevin Lingerfelt	eebc612d52	Add install flag for sending tls identity info to proxies (#1055 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-04 16:55:06 -07:00
Kevin Lingerfelt	ec2433e9bd	Update controller to use 'tls' metric label (#1044 ) * Update controller to use 'tls' metric label * Fix meshed column formatter Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-06-01 16:44:33 -07:00
Sacha Froment	84781c9c74	conduit inject: Add flag to set proxy bind timeout (#865 ) * conduit inject: Add flag to set proxy bind timeout (#863) * fix test * fix flag to get it working with #909 * Add time parsing * Use the variable to set the default value Signed-off-by: Sacha Froment <sfroment42@gmail.com>	2018-05-29 11:14:29 -07:00
Risha Mars	d333a7d861	Add a secured label to the CLI tap responses (#996 ) Adds secured=yes/no to the conduit tap responses. This assumes a `meshed` label is returned by the proxy.	2018-05-25 11:21:38 -07:00
Risha Mars	ffabdefc6c	Add queries to prometheus to determine number of fully meshed requests (#983 ) - Update the `response_total` prometheus query of the StatSummary endpoint to also break queries out by a `meshed` label. - Add a 'Secured' column to the web UI/CLI stat displays, which indicate the percentage of traffic starting and ending in the mesh This meshed label is used in the CLI/Web UI to display a column of the percentage of traffic that starts/ends in the mesh. (Which is a proxy indicator for whether that traffic is 'secured' when we add TLS by default for intra mesh requests). The `meshed` label is not yet added anywhere, so until it is supplied by the proxy, all traffic will show up as 0% secured in the web/CLI.	2018-05-24 11:05:09 -07:00
Haiwei Liu	8c98cde82b	change init image to root options, for install and inject to use (#1001 ) Signed-off-by: Haiwei Liu <carllhw@gmail.com>	2018-05-24 10:21:16 -07:00
Andrew Seigner	8a1a3b31d4	Fix non-default proxy-api port (#979 ) Running `conduit install --api-port xxx` where xxx != 8086 would yield a broken install. Fix the install command to correctly propagate the `api-port` flag, setting it as the serve address in the proxy-api container. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-05-22 10:34:25 -07:00
Kevin Lingerfelt	2baeaacbc8	Remove package-scoped vars in cmd package (#975 ) * Remove package-scoped vars in cmd package * Run gofmt on all cmd package files * Re-add missing Args setting on check command Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-05-21 18:15:39 -07:00
Kevin Lingerfelt	36ec391dbe	Go: update k8s dependencies to 1.10.2 (#962 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-05-17 15:46:58 -07:00
Risha Mars	b8dc83f9d2	Modify the Stat API to handle requests for resource type "all" (#928 ) Allow the Stat endpoint in the public-api to accept requests for resourceType "all". Currently, this queries Pods, Deployments, RCs and Services, but can be modified to query other resources as well. Both the CLI and web endpoints now work if you set resourceType to all. e.g. `conduit stat all`	2018-05-11 14:35:37 -07:00
Kevin Lingerfelt	4e8e1eb84d	CLI: Fix validation for service stats (#935 ) * CLI: Fix validation for service stats * Address review feedback Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-05-11 10:28:49 -07:00
Oliver Gould	a786089fd6	docker: Cache versionless builds before building versioned go binaries (#921 ) The way that git-related version information is linked into go binaries busts Docker's cache such that every commit causes all binaries to rebuilt. In order to ameliorate this, we can build each binary once without version information first so that its artifacts are cached. When Go sources are not changed and only the version information changes, builds are 4.3x faster than before (from 5+ minutes to <90s). On `master` Branch off of master and build (mostly cached): ``` :; time DOCKER_TRACE=1 bin/docker-build ... DOCKER_TRACE=1 bin/docker-build 9.10s user 6.30s system 5% cpu 4:26.47 total ``` Rebuild without changing anything (highly cached): ``` :; time DOCKER_TRACE=1 bin/docker-build ... DOCKER_TRACE=1 bin/docker-build 9.23s user 6.04s system 47% cpu 32.017 total ``` Update only the git sha and rebuild: ``` :; git ci -am 'bump it' --allow-empty [ver/eg 2749eb3] bump it :; time DOCKER_TRACE=1 bin/docker-build ... DOCKER_TRACE=1 bin/docker-build 8.55s user 6.08s system 4% cpu 5:22.25 total ``` On this branch: Rebuild without changing anything (highly cached): ``` :; time DOCKER_TRACE=1 bin/docker-build ... DOCKER_TRACE=1 bin/docker-build 8.94s user 5.97s system 46% cpu 32.257 total ``` Update only the git sha and rebuild: ``` :; git ci -am 'bump it' --allow-empty [ver/go-docker-cache-versionless 77a80b5] bump it :; time DOCKER_TRACE=1 bin/docker-build ... DOCKER_TRACE=1 bin/docker-build-cli-bin 2.02s user 1.34s system 9% cpu 34.144 total ```	2018-05-10 10:22:09 -07:00
Thomas Rampelberg	25d8b22b5c	Adding statefulsets to inject. Fixes #907 (#910 )	2018-05-10 09:00:36 -05:00
Thomas Rampelberg	e96c2e1135	Made inject aware of the List type. (#886 )	2018-05-10 08:08:55 -05:00
Andrew Seigner	1275b1ae89	Introduce Grafana, K8s, and Prom dashboards (#904 ) Grafana provides default dashboards for Prometheus and Grafana health. The community also provides Kubernetes-specific dashboards. Conduit was not taking advantage of these. Introduce new Grafana dashboards focused on Grafana, Kubernetes, and Prometheus health. Tag all Conduit dashboards for easier UI navigation. Also fix layout in Conduit Health dashboard. Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-05-08 23:11:43 +02:00
Risha Mars	f94856e489	Modify the Stat endpoint to also return the number of failed conduit pods (#895 ) * Modify the Stat endpoint to also return the count of failed pods * Add comments explaining pod count stats * Rename total pod count to running pod count This is to support the service mesh overview page, as I'd like to include an indicator of failed pods there.	2018-05-08 10:35:21 -07:00
Andrew Seigner	97bf4fcdf2	Release Notes for 0.4.1 release. (#839 ) Also update Getting Started and Debugging docs to reflect changes in `Tap` and `Stat`. Fixes #838 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-26 13:32:41 -07:00
Brian Smith	c5d2dab8bd	Remove special support for ExternalName services (#764 ) After this was implemented we found that ExternalName services are represented in DNS as CNAMEs, which means that the proxy's DNS fallback logic can be used instead of doing DNS in the control plane. Besides simplifying the controller, this will also increase fidelity with the proxied pods' DNS configuration (improve transparency). Signed-off-by: Brian Smith <brian@briansmith.org>	2018-04-25 11:53:33 -10:00
Andrew Seigner	dce31b888f	Deprecate Tap, rename TapByResource to Tap (#844 ) The `conduit tap` command is now deprecated. Replace `conduit tap` with `connduit tapByResource`. Rename tapByResource to tap. The underlying protobuf for tap remains, the tap gRPC endpoint now returns Unimplemented. Fixes #804 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-25 12:24:46 -07:00
Andrew Seigner	640570cd6b	Make TapByResource output destination pod (#837 ) The TapByResource command now has access to destination labels from the proxy, but was not outputting them on the cli. Modify the TapByResource output to print the destination pod label, rather than the ip, when available. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-24 18:58:40 -07:00
Andrew Seigner	a0a9a42e23	Implement Public API and Tap on top of Lister (#835 ) public-api and and tap were both using their own implementations of the Kubernetes Informer/Lister APIs. This change factors out all Informer/Lister usage into the Lister module. This also introduces a new `Lister.GetObjects` method. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-24 18:10:48 -07:00
Andrew Seigner	baf4ea1a5a	Implement TapByResource in Tap Service (#827 ) The TapByResource endpoint was previously a stub. Implement end-to-end tapByResource functionality, with support for specifying any kubernetes resource(s) as target and destination. Fixes #803, #49 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-23 16:13:26 -07:00
Andrew Seigner	39eccb09e2	cli: standardize kubernetes resource parsing (#830 ) The Tap command leveraged new cli parsing code, enabling Kubernetes resources specified as `(TYPE [NAME] \| TYPE/NAME)`. The Stat command did not use this. Modify the Stat command to use the same cli flag parsing code as Tap. Remove the to/from-resource flags from Stat. Fixes #792 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-23 15:17:42 -07:00
Andrew Seigner	1f32f130de	Upgrade Prometheus from 2.1.0 to 2.2.1 (#816 ) There have been a number of performance improvements and bug fixes since v2.1.0. Bump our Prometheus container to v2.2.1. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-19 18:00:53 -07:00
Andrew Seigner	79bdc638b3	Service support in stat command (#809 ) The `stat` command did not support `service` as a resource type. This change adds `service` support to the `stat` command. Specifically: - as a destination resource on `--to` commands - as a target resource on `--from` commands Fixes #805 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-19 16:51:20 -07:00
Andrew Seigner	293e00bc3e	Introduce tapByResource cli command (#802 ) The existing `tap` command is being deprecated. Introduce a `tapByResource` cli command. It supports tapping a Kubernetes resource or collection of resources, optionally filtered by outbound resources. This command will eventually replace `tap`. Part of #778 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-19 14:44:23 -07:00
Kevin Lingerfelt	653dc6bfaa	Add replication controller stats in CLI (#794 ) * Add replication controller stats in CLI * Fix pod status in stat summary tests Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-18 18:12:14 -07:00
Oliver Gould	06dd8d90ee	Introduce the TapByResource API (#778 ) This changes the public api to have a new rpc type, `TapByResource`. This api supersedes the Tap api. `TapByResource` is richer, more closely reflecting the proxy's capabilities. The proxy's Tap api is extended to select over destination labels, corresponding with those returned by the Destination api. Now both `Tap` and `TapByResource`'s responses may include destination labels. This change avoids breaking backwards compatibility by: * introducing the new `TapByResource` rpc type, opting not to change Tap * extending the proxy's Match type with a new, optional, `destination_label` field. * `TapEvent` is extended with a new, optional, `destination_meta`.	2018-04-18 15:37:07 -07:00
Andrew Seigner	1e4ac8fda8	Destination service provides pod-template-hash (#784 ) The Destination service does not provide ReplicaSet information to the proxy. The `pod-template-hash` label approximates selecting over all pods in a ReplicaSet or ReplicationController. Modify the Destination service to provide this label to the proxy. Relates to #508 and #741 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-18 14:41:27 -07:00
Kevin Lingerfelt	71a51afb40	Expose pod stats in CLI, web UI, and Grafana (#788 ) * Expose pod stats in CLI, web UI, and Grafana * Fix js api helpers test * Add outbound traffic stats to pod dashboard Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-18 11:26:47 -07:00
Andrew Seigner	727521f914	Permit arbitrary time windows in public-api (#774 ) The public-api previously only permitted 4 hard-coded time windows: 10s, 1m, 10m, 1h. This was primarily a relic of the recently removed telemetry system. Modify the public-api to validate the time string, but allow for any window size, which is then passed through to Prometheus. Fixes #686 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-16 17:37:17 -07:00
Kevin Lingerfelt	11a4359e9a	Misc cleanup following the telemetry rewrite (#771 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-16 15:51:07 -07:00
Oliver Gould	800cefdb77	Skip the proxy on the metrics port (#770 ) When prometheus queries the proxy for data, these requests are reported as inbound traffic to the pod. This leads to misleading stats when a pod otherwise receives little/no traffic. In order to prevent these requests being proxied, the metrics port is now added to the default inbound skip-ports list (as is already case for the tap server). Fixes #769	2018-04-16 11:54:58 -07:00
Andrew Seigner	77fb6d3709	Add namespace as a resource type in public-api (#760 ) * Add namespace as a resource type in public-api The cli and public-api only supported deployments as a resource type. This change adds support for namespace as a resource type in the cli and public-api. This also change includes: - cli statsummary now prints `-`'s when objects are not in the mesh - cli statsummary prints `No resources found.` when applicable - removed `out-` from cli statsummary flags, and analagous proto changes - switched public-api to use native prometheus label types - misc error handling and logging fixes Part of #627 Signed-off-by: Andrew Seigner <siggy@buoyant.io> * Refactor filter and groupby label formulation Signed-off-by: Kevin Lingerfelt <kl@buoyant.io> * Rename stat_summary.go to stat.go in cli Signed-off-by: Kevin Lingerfelt <kl@buoyant.io> * Update rbac privileges for namespace stats Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-13 16:53:01 -07:00
Oliver Gould	cc44db054f	Remove NODE_NAME and POD_NAME env usage (#758 ) * proxy: Remove pod_name and node_name * cli: Do not inject POD_NAME and NODE_NAME env vars	2018-04-13 13:09:51 -07:00
Andrew Seigner	21886760c6	Use apps/v1beta2 for Kubernetes 1.8 compatibility (#762 ) Conduit was relying on apps/v1 to Deployment and ReplicaSet APIs. apps/v1 is not available on Kubernetes 1.8. This prevented the public-api from starting. Switch Conduit to use apps/v1beta2. Also increase the Kubernetes API cache sync timeout from 10 to 60 seconds, as it was taking 11 seconds on a test cluster. Fixes #761 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-13 12:08:16 -07:00
Kevin Lingerfelt	fb15fe7c1a	Remove the telemetry service (#757 ) * Remove the telemetry service The telemetry service is no longer needed, now that prometheus scrapes metrics directly from proxies, and the public-api talks directly to prometheus. In this branch I'm removing the service itself as well as all of the telemetry protobuf, and updating the conduit install command to no longer install the service. I'm also removing the old version of the stat command, which required the telemetry service, and renaming the statsummary command to stat. * Fix time window tests * Remove deprecated controller scrape config Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-13 11:21:29 -07:00
Kevin Lingerfelt	47caf1ca07	Add --all-namespaces flag to CLI statsummary command (#745 ) * Add --all-namespaces flag to CLI statsummary command * Fix statsummary output formatting Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-11 16:40:25 -07:00
Andrew Seigner	259fdcd134	Add latency stats in new stat summary endpoint (#737 ) The new StatSummary endpoint was only providing request volume and successs rate information. Add support for retrieving latency stats via StatSummary. Also make all prometheus calls in parallel, and implement kubernetes test fixtures. Fixes #681 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-11 11:58:32 -07:00
Kevin Lingerfelt	91c359e612	Switch public API to use cached k8s resources (#724 ) * Switch public API to use cached k8s resources * Move shared informer code to separate goroutine * Fix spelling issue Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-10 11:39:31 -07:00
Andrew Seigner	716b392231	Move StatSummary logic into grpc server (#717 ) The StatSummary logic was implemented as a method on http_server. Move the StatSummary logic into grpc_server, for consistency with the other endpoints. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-06 16:46:15 -07:00
Kevin Lingerfelt	baa4d10c2f	CLI: change conduit namespace shorthand flag to -c (#714 ) * CLI: change conduit namespace shorthand flag to -c All of the conduit CLI subcommands accept a --conduit-namespace flag, indicating the namespace where conduit is running. Some of the subcommands also provide a --namespace flag, indicating the kubernetes namespace where a user's application code is running. To prevent confusion, I'm changing the shorthand flag for the conduit namespace to -c, and using the -n shorthand when referring to user namespaces. As part of this change I've also standardized the capitalization of all of our command line flags, removed the -r shorthand for the install --registry flag, and made the global --kubeconfig and --api-addr flags apply to all subcommands. * Switch flag descriptions from lowercase to Capital Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-04-06 14:47:31 -07:00
Andrew Seigner	50c323c617	Use canonical k8s names, fix prom labels (#702 ) The new statsummary command accepted friendly k8s names, which worked for k8s queries, but Prometheus requires a specific key. Modify the statsummary query to map friendly k8s names to canonical k8s names when constructing the query. Then during the query, map the canonical k8s name to a specific Prometheus label. Fixes #695 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-06 12:34:54 -07:00
Risha Mars	2f5b5ea5f2	Start implementing conduit stat summary endpoint (#671 ) Start implementing new conduit stat summary endpoint. Changes the public-api to call prometheus directly instead of the telemetry service. Wired through to `api/stat` on the web server, as well as `conduit statsummary` on the CLI. Works for deployments only. Current implementation just retrieves requests and mesh/total pod count (so latency stats are always 0). Uses API defined in #663 Example queries the stat endpoint will eventually satisfy in #627 This branch includes commits from @klingerf * run ./bin/dep ensure * run ./bin/update-go-deps-shas	2018-04-05 17:05:06 -07:00
Andrew Seigner	28d5007cdf	Harmonize Prometheus label usage (#690 ) The Destination service used slightly different labels than the telemetry pipeline expected, specifically, prefixed with `k8s_`. Make all Prometheus labels consistent by dropping `k8s_`. Also rename `pod_name` to `pod` for consistency with `deployement`, etc. Also update and reorganize `proxy-metrics.md` to reflect new labelling. Fixes #655 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-05 15:09:06 -07:00
Andrew Seigner	9508e11b45	Build conduit-specific Grafana Docker image (#679 ) Using a vanilla Grafana Docker image as part of `conduit install` avoided maintaining a conduit-specific Grafana Docker image, but made packaging dashboard json files cumbersome. Roll our own Grafana Docker image, that includes conduit-specific dashboard json files. This significantly decreases the `conduit install` output size, and enables dashboard integration in the docker-compose environment. Fixes #567 Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-05 14:20:05 -07:00
Andrew Seigner	ee042e1943	Rename grafana viz to top-line (#666 ) The primary Grafana dashboard was named 'viz' from a prototype. Rename 'viz' to 'Top Line'. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-02 18:10:35 -07:00
Brian Smith	df9ead9c36	Use Go 1.10.1 to build all Go code. (#650 ) Go 1.10.1 is a security release. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-04-02 14:58:30 -10:00
Andrew Seigner	bf721466e3	Filter out conduit controller pods from Grafana (#657 ) The Grafana dashboards were displaying all proxy-enabled pods, including conduit controller pods. In the old telemetry pipeline filtering these out required knowledge of the controller's namespace, which the dashboards are agnostic to. This change leverages the new `conduit_io_control_plane_component` prometheus label to filter out proxy-enabled controller components. Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-04-02 17:56:12 -07:00
Andrew Seigner	8fe742e2de	Update Grafana dashboards to use new proxy metrics (#637 ) The Top-line and Deployment Grafana dashboards relied on the soon-to-be-removed telemetry pipeline metrics. Update the Grafana dashboards to query for the new, proxy-based metrics. Grafana dashboard layouts have not changed. Depends on #635 to render metrics. Part of #420. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-29 13:00:01 -07:00
Andrew Seigner	666c83e963	Add pod_name to Prometheus labels (#649 ) Previously we were using the instance label to uniquely identify a pod. This meant that getting stats by pod name would require extra queries to Kubernetes to map pod name to instance. This change adds a pod_name label to metrics at collection time. This should not affect cardinality as pod_name is invariant with respect to instance. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-29 11:07:35 -07:00
Deshi Xiao	732f3c1565	fix grafana CrashLoopBackoff on image 5.0.3 (#646 ) this is a known issue with grafana in k8s. grafana/grafana:5.0.4 was just released today. update the repo from 5.0.3 to 5.0.4 fixed issues #582 Signed-off-by: Deshi Xiao <xiaods@gmail.com>	2018-03-29 09:36:11 -07:00
Oliver Gould	6e435754a1	Improve CLI docker caching (#612 ) Currently, the CLI docker image copies the entire `controller` directory, though the CLI only requires a few of its subdirectories. This causes the CLI's docker cache to be needlessly invalidated when, for instance, a service implementation changes. By restricting the copied directories to `controller/{api,public,util}`, build caching is improved.	2018-03-29 09:29:06 -07:00
Kevin Lingerfelt	59c75a73a9	Add tests/utils/scripts for running integration tests (#608 ) * Add tests/utils/scripts for running integration tests Add a suite of integration tests in the `test/` directory, as well as utilities for testing in the `testutil/` directory. You can use the `bin/test-run` script to run the full suite of tests, and the `bin/test-cleanup` script to cleanup after the tests. The test/README.md file has more information about running tests. @pcalcado, @franziskagoltz, and @rmars also contributed to this change. * Create TEST.md file at the root of the repo * Update based on review feedback * Relax external service IP timeout for GKE * Update TEST.md with more info about different types of test runs * More updates to TEST.md based on review feedback Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-03-27 15:06:55 -07:00
Andrew Seigner	fe35509406	Clean up Prometheus labels scraped from proxy (#633 ) The Prometheus scrape config collects from Conduit proxies, and maps Kubernetes labels to Prometheus labels, appending "k8s_". This change keeps the resultant Prometheus labels consistent with their source Kubernetes labels. For example: "deployment" and "pod_template_hash". Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-27 15:01:08 -07:00
Andrew Seigner	291d8e97ab	Move injected data from env var to k8s labels (#605 ) The inject code detects the object it is being injected into, and writes self-identifying information into the CONDUIT_PROMETHEUS_LABELS environment variable, so that conduit-proxy may read this information and report it to Prometheus at collection time. This change puts the self-identifying information directly into Kubernetes labels, which Prometheus already collects, removing the need for conduit-proxy to be aware of this information. The resulting label in Prometheus is recorded in the form `k8s_deployment`. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-23 16:11:34 -07:00
Andrew Seigner	fb1d6a5c66	Introduce Conduit Health dashboard (#591 ) In addition to dashboards display service health, we need a dashboard to display health of the Conduit service mesh itself. This change introduces a conduit-health dashboard. It currently only displays health metrics for the control plane components. Proxy health will come later. Fixes #502 Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-22 15:16:03 -07:00
Alex Leong	d50550515e	Add the proxy pod owner as a Prometheus label (#448 ) Update the inject command to set a CONDUIT_PROMETHEUS_LABELS proxy environment variable with the name of the pod spec that the proxy is injected into. This will later be used as a label value when the proxy is exposing metrics. Fixes: #426 Signed-off-by: Alex Leong <alex@buoyant.io>	2018-03-22 15:10:51 -07:00
Andrew Seigner	c03508ba8c	Update Prometheus to scrape data and control plane (#583 ) The existing telemetry pipeline relies on Prometheus scraping the Telemetry service, which will soon be removed. This change configures Prometheus to scrape the conduit proxies directly for telemetry data, and the control plane components for control-plane health information. This affects the output of both conduit install and conduit inject. Fixes #428, #501 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-22 13:58:11 -07:00
Andrew Seigner	680bf6211a	Add Grafana support to conduit dashboard command (#590 ) The existing `conduit dashboard` command supported opening the conduit dashboard, or displaying the conduit dashboard URL, via a `url` boolean flag. Replace the `url` boolean flag with a `show` string flag, with three modes: `conduit dashboard --show conduit`: default, open conduit dashboard `conduit dashboard --show grafana`: open grafana dashboard `conduit dashboard --show url`: display dashboard URLs Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-20 18:07:30 -07:00
Andy Hume	e6286e1bdf	cli: ensure check command has 80-character output (#587 ) Successful `conduit check` commands now take into account `[ok]` and `\n` tokens when constraining line length. Fixes #554 Signed-off-by: Andy Hume <andyhume@gmail.com>	2018-03-20 13:55:19 -07:00
Andrew Seigner	3ca8e84eec	Add Top Line and Deployment Grafana dashboards (#562 ) Existing Grafana configuration contained no dashboards, just a skeleton for testing. Introduce two Grafana dashboards: 1) Top Line: Overall health of all Conduit-enabled services 2) Deployment: Health of a specific conduit-enabled deployment Fixes #500 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-20 10:22:30 -07:00
Brian Smith	e7c4a9d4b9	Remove the cli docker image (#579 ) This image isn't used. It references its base image using the `latest` tag, which is wrong; it should have been using the tag that the base image was built with. It is likely that the last few iterations of this image that we've published have wrong and useless contents. With that in mind, just remove the image. Fixes #578. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-03-16 14:22:46 -10:00
Alex Leong	9eb084c99d	Most controller listeners should only bind on localhost (#494 ) * Most controller listeners should only bind on localhost * Use default listening addresses in controller components * Review feedback * Revert test_helper change * Revert use of absolute domains Signed-off-by: Alex Leong <alex@buoyant.io>	2018-03-12 11:32:20 -07:00
Brian Smith	649e784d9c	Simplify cluster zone suffix handling in the proxy (#528 ) * Temporarily stop trying to support configurable zones in the proxy. None of the zone configuration is tested and lots of things assume the cluster zone is `cluster.local`. Further, how exactly the proxy will actually learn the cluster zone hasn't been decided yet. Just hard-code the zone as "cluster.local" in the proxy until configurable zones are fully implemented and tested to be working correctly. Signed-off-by: Brian Smith <brian@briansmith.org> * Remove the CONDUIT_PROXY_DESTINATIONS_AUTOCOMPLETE_FQDN setting The way that Kubernetes configures DNS search suffixes has some negative consequences as some names like "example.com" are ambiguous: depending on whether there is a service "example" in the "com" namespace, "example.com" may refer to an external service or an internal service, and this can fluctuate over time. In recognition of that we added the CONDUIT_PROXY_DESTINATIONS_AUTOCOMPLETE_FQDN setting, thinking this would be part of a solution for users to opt out of the unfortunate behavior if their applications didn't depend on the DNS search suffix feature. It turns out similar effects can be acheived using a custom dnsConfig, starting in Kubernetes 1.10 when dnsConfig reaches the beta stability level. Now any CONDUIT_PROXY_DESTINATIONS_AUTOCOMPLETE_FQDN-based seems duplicative. Further, attempting to support it optionally made the code complex and hard to read. Therefore, let's just remove it. If/when somebody actually requests this functionality then we can add it back, if dnsConfig isn't a valid alternative for them. Signed-off-by: Brian Smith <brian@briansmith.org> * Further hard-code "cluster.local" as the zone, temporarily. Addresses review feedback. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-03-07 14:30:13 -10:00
Dennis Adjei-Baah	ad42f2f8ab	Retry k8s watch endpoints on error (#510 ) Shortly after conduit is installed in k8s environment. The control plane component that establishes a watch endpoint with k8s run in to networking issues during proxy initialization. During failure, each watcher fails to retry its connection to k8s watch endpoint which leads to timeouts and eventually, multiple controller pod restarts. This PR adds retry logic to each "watch" enabled package. fixes #478 Signed-off-by: Dennis Adjei-Baah <dennis@buoyant.io>	2018-03-07 13:40:43 -08:00
Brian Smith	0d4ab39ce7	Revert "Make absolute names truly absolute. (#525 )" (#533 ) This reverts commit `517616a166`. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-03-07 10:57:10 -10:00
Brian Smith	517616a166	Make absolute names truly absolute. (#525 ) Kubernetes will do multiple DNS lookups for a name like `proxy-api.conduit.svc.cluster.local` based on the default search settings in /etc/resolv.conf for each container: 1. proxy-api.conduit.svc.cluster.local.conduit.svc.cluster.local. IN A 2. proxy-api.conduit.svc.cluster.local.svc.cluster.local. IN A 3. proxy-api.conduit.svc.cluster.local.cluster.local. IN A 4. proxy-api.conduit.svc.cluster.local. IN A We do not need or want this search to be done, so avoid it by making each name absolute by appending a period so that the first three DNS queries are skipped for each name. The case for `localhost` is even worse because we expect that `localhost` will always resolve to 127.0.0.1 and/or ::1, but this is not guaranteed if the default search is done: 1. localhost.conduit.svc.cluster.local. IN A 2. localhost.svc.cluster.local. IN A 3. localhost.cluster.local. IN A 4. localhost. IN A Avoid these unnecessary DNS queries by making each name absolute, so that the first three DNS queries are skipped for each name. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-03-07 09:46:03 -10:00
Kevin Lingerfelt	47fc2eae20	Set -logtostderr flag on controller components (#524 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-03-07 10:18:15 -08:00
Andrew Seigner	a065174688	Disable Grafana update check (#521 ) Grafana by default calls out to grafana.com to check for updates. As user's of Conduit do not have direct control over updating Grafana directly, this update check is not needed. Disable Grafana's update check via grafana.ini. This is also a workaround for #155, root cause of #519. Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-03-06 16:14:44 -08:00
Kevin Lingerfelt	d6bd17425a	Add --expected-version flag for conduit check command (#497 ) * Add --expected-version flag for conduit check command Signed-off-by: Kevin Lingerfelt <kl@buoyant.io> * Update build instructions Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-03-06 11:32:14 -08:00
Dennis Adjei-Baah	5a4c5aa683	Exclude telemetry generated by the control plane when requesting depl… (#493 ) When the conduit proxy is injected into the controller pod, we observe controller pod proxy stats show up as an "outbound" deployment for an unrelated upstream deployment. This may cause confusion when monitoring deployments in the service mesh. This PR filters out this "misleading" stat in the public api whenever the dashboard requests metric information for a specific deployment. * exclude telemetry generated by the control plane when requesting deployment metrics fixes #370 Signed-off-by: Dennis Adjei-Baah <dennis@buoyant.io>	2018-03-05 17:58:08 -08:00
Brian Smith	4c9b9c0f68	Install: Don't install buoyantio/kubectl into the prometheus pod. (#509 ) In the initial review for this code (preceding the creation of the runconduit/conduit repository), it was noted that this container is not actually used, so this is actually dead code. Further, this container actualy causes a minor problem, as it doesn't implement any retry logic, thus it will sometimes often cause errors to be logged. See https://github.com/runconduit/conduit/issues/496#issuecomment-370105328. Further, this is a "buoyantio/" branded container. IF we actually need such a container then it should be a Conduit-branded container. See https://github.com/runconduit/conduit/issues/478 for additional context. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-03-05 08:59:14 -10:00
Igor Zibarev	0f6db6efc0	cli: refactor k8s config to support $KUBECONFIG with multiple paths (#482 ) Kubernetes $KUBECONFIG environment variable is a list of paths to configuration files, but conduit assumes that it is a single path. Changes in this commit introduce a straightforward way to discover and load config file(s). Complete list of changes: - Use k8s.io/client-go/tools/clientcmd to deal with kubernetes configuration file - rename k8s API and k8s proxy constructors to get rid of redundancy - remove shell package as it is not needed anymore Signed-off-by: Igor Zibarev <zibarev.i@gmail.com>	2018-02-28 12:13:09 -08:00
Andrew Seigner	d50c8b4ac8	Add Grafana to conduit install (#444 ) `conduit install` deploys prometheus, but lacks a general-purpose way to visualize that data. This change adds a Grafana container to the `conduit install` command. It includes two sample dashboards, viz and health, in their own respective source files. Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-02-28 11:36:21 -08:00
Andy Hume	1e611c21c6	cli: add check for latest version of cli and control plane (#460 ) As part of `conduit check` command, warn the user if they are running an outdated version of the cli client or the control plane components. Fixes #314 Signed-off-by: Andy Hume <andyhume@gmail.com>	2018-02-27 16:11:38 -08:00
Dennis Adjei-Baah	893bacf8d6	Make prometheus URL in config fully qualified DNS name (#443 ) The telemetry service in the controller pod uses a non-fully qualified URL to connect to the prometheus pod in the control plane. This PR changes the URL the telemetry's prometheus URL to be fully qualified to be consistent with other URLs in the control plane. This change was tested in minikube. The logs report no errors and looking at the prometheus dashboard shows that stats are being recorded from all conduit proxies. fixes #414 Signed-off-by: Dennis Adjei-Baah dennis@buoyant.io	2018-02-26 09:40:31 -08:00
Brian Smith	34cf79a3e6	Add a test of the actual default output of `conduit install`. (#376 ) Refactor `conduit install` test into a data-driven test. Then add a test of the actual default output of `conduit install`. This test is useful to make it clear when we change the default settings of `conduit install`. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-02-23 13:27:36 -10:00
Brian Smith	78ebd5e340	Base control plane Docker images on scratch instead of base. (#368 ) The control plane is proxied through the Conduit proxy. The Conduit proxy is based on the base image, and the control plane containers and the proxy share a networking namespace. This means we don't need the extra base utilities in the controller images since we can use the utilties in the proxy image. This is a step towards building the initial no-networking Conduit CA pod. Since the Conduit CA will not do any networking of its own, we networking debugging utilties are not helpful for it. They are actually an unnecessary risk because they could facilitate the exfiltration of the private key of the CA. (The Conduit CA pod won't have the Conduit Proxy injected into it either.) This also simplifies & slightly speeds up the building of the controller images. This is a stepping stone towards being able to build the controller images without `docker build` to improve build times. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-02-23 13:03:19 -10:00
Brian Smith	86bb65a148	Remove potentially-conflicting `app` labels in control plane (#373 ) The `app` label should be reserved for end-user applications and we shouldn't use it ourselves. We already have a Conduit-specific label that is is prefixed with the `conduit.io/` prefix to avoid naming collisions with users' labels, so just use that one instead. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-02-23 12:43:55 -10:00
Andrew Seigner	83f9c391bb	Move install template to its own file (#423 ) The template used by `conduit install` was hard-coded in install.go. This change moves the template into its own file, in anticipation of increasing the template's size and complexity. Part of #420 Signed-off-by: Andrew Seigner <siggy@buoyant.io>	2018-02-23 14:15:31 -08:00
Dennis Adjei-Baah	f66ec6414c	Inject the conduit proxy into controller pod during conduit install (#365 ) In order to take advantage of the benefits the conduit proxy gives to deployments, this PR injects the conduit proxy into the control plane pod. This helps us lay the groundwork for future work such as TLS, control plane observability etc. Fixes #311 Signed-off-by: Dennis Adjei-Baah <dennis@buoyant.io>	2018-02-23 13:55:46 -08:00
Brian Smith	cf3c8cd7bc	Use Go 1.10.0 to build Go components. (#408 ) * Use Go 1.10.0 to build Go components. Take advantage of the new build cache in Go 1.10. Future work on improving build performance will utilize the build cache further. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-02-21 14:31:29 -10:00
Brian Smith	e6aad57766	Remove temporary files generated by dep in go-deps image. (#407 ) Previously Dockerfile-go-deps was converted from a multi-stage Dockefile to a single-stage Dockerfile in anticipation of enabling efficient use of `--cache-from` in CI. However, that resulted in the image ballooning in size because it contained the Git repo for every package downloaded by `dep ensure`. Bring the image back down to the proper size by removing the temporary files created. Signed-off-by: Brian Smith <brian@briansmith.org>	2018-02-21 13:06:24 -10:00
Kevin Lingerfelt	c579a8fe8d	Improve get/stat/tap help text by way of examples (#401 ) Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-02-20 19:27:42 -08:00
Kevin Lingerfelt	8db7115420	Update go-run to set version equal to root-tag (#393 ) * Update go-run to set version equal to root-tag * Fix inject tests for undefined version change * Pass inject version explitictly as arg Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-02-20 12:25:55 -08:00
Kevin Lingerfelt	f48555d3cc	Remove kubectl dependency, validate k8s server version via api (#396 ) * Remove kubectl dependency, validate k8s server version via api Signed-off-by: Kevin Lingerfelt <kl@buoyant.io> * Remove unused MockKubectl Signed-off-by: Kevin Lingerfelt <kl@buoyant.io> * Remame kubectl.go to version.go Signed-off-by: Kevin Lingerfelt <kl@buoyant.io>	2018-02-20 12:14:11 -08:00
Dennis Adjei-Baah	9af3783555	Print error message only when invalid YAML file is used with inject command (#389 ) When the `inject` command is used on a YAML file that is invalid, it prints out an invalid YAML file with the injected proxy. This may give a false indication to the user that the inject was successful even though the inject command prints out an error message further down the terminal window. This PR fixes #303 and contains a test input and output file that indicates what should be shown. This PR also fixes #390. Signed-off-by: Dennis Adjei-Baah <dennis@buoyant.io>	2018-02-20 11:59:41 -08:00

1 2 3 4 5 ...

315 Commits